How a Firebase and Cloud Run Test Ran Up a $72,000 Bill
- Platform
- Firebase / Google Cloud
- When
- March 27, 2020
- Bill
- $72,000
- Outcome
- Google waived the bill as a one-time gesture, per the founder
In March 2020, Milkie Way, a small startup, ran an internal test of a web scraper on Google Cloud Run with Firebase's Cloud Firestore as its database. The scraper kept sending new requests to itself, and within hours the project had a bill of a little under $72,000, against a $7 budget. Founder Sudeep Chauhan described the incident in a two-part post-mortem published in December 2020, and Google later waived the bill.
What happened
- Early 2020: The team was building Announce-AI, a web scraper. Cloud Functions timed out after about 9 minutes, so they tried Cloud Run.
- Setup: Chauhan created a new project, set a $7 Cloud Billing budget, and left the Firebase project on the free Spark plan. The company card had a $100 spending limit.
- Test day: The team deployed the code, sent a few manual requests, checked logs and billing for a few minutes, saw nothing unusual and moved on.
- Friday, March 27, 2020: After a nap the next afternoon, Chauhan found three emails: the Firebase project had been upgraded to a paid plan, the budget was exceeded, and the card was declined. The billing console showed about $5,000, then $15,000 five minutes later and $25,000 after 20 minutes. After two hours it settled just under $72,000. The team disabled billing and shut down services. Because every project used the same card, Google suspended all of them, three days before Announce was due to launch.
- March 28 onward: Chauhan contacted law firms and wrote an incident report for Google, which took about 10 days to respond.
- December 8, 2020: The post-mortem was published and was widely discussed on Hacker News; The Register covered it on December 10.
Why the bill got so big
To get around the timeout, each Cloud Run request scraped a single page and then sent a new POST request back to the same service for every link it found. There was no stop condition and no check for URLs already visited. When a page linked back to an earlier one, the requests looped, and because each page produced several new requests, the work multiplied instead of just repeating.
The service ran with what Chauhan says were Cloud Run's defaults at the time: up to 1,000 instances and 80 concurrent requests per instance. Each instance read from and wrote to Firestore every few milliseconds. At the peak, Firestore served about one billion reads per minute.
- Firestore: 116 billion reads and 33 million writes. At the $0.06 per 100,000 reads he cites, the reads alone came to $69,600, most of the bill.
- Cloud Run: When logging stopped, the team assumed the requests had died, but work continued in the background and the services were never deleted. Chauhan reports 16,022 hours of compute in 24 hours.
Three things hid the problem. Because the Google Cloud project had billing attached for Cloud Run, the Firebase project was upgraded from Spark to paid automatically. Billing data arrived about a day late. And the Firebase console still showed about 42,000 reads and writes for the month after the bill had arrived.
The post-mortem is not consistent on timing. Part 1 says the money was spent within a few hours, Part 2 says the 116 billion reads took less than an hour, and the Cloud Run figure covers 24 hours. Chauhan estimated that with a maximum of 2 instances instead of 1,000, the bill would have been about $144.
How it ended
After reviewing the team's incident report, along with consultations and internal discussion, Google "let go of our bill as a one-time gesture," Chauhan wrote. The Register's coverage includes no statement from Google, so the founder's account is the only public record of the outcome. Announce launched in late November 2020, about seven months late.
How to protect yourself on Firebase and Google Cloud
- Know what budgets do. A standard budget only sends alerts and does not cap usage or spending. Google notes that cost reporting lags usage and that the first notification after you create a budget can take several hours, so set the amount below what you can afford to lose.
- Use spend caps where they exist. Spend cap budgets, currently in Preview, pause a service for the rest of the month once it reaches 100% of its budget. On Google Cloud they cover Cloud Run, Cloud Run functions and the Gemini APIs. In Firebase they cover Cloud Functions, App Hosting, Extensions and AI Logic. Firestore is not on either list, and enforcement can lag by several minutes, with any overage billed as normal. In a case like this one, a cap on Cloud Run could have paused the service that was generating the reads.
- Disable billing automatically if an outage is acceptable. Google documents a setup where a budget publishes to Pub/Sub and a function removes billing from the project. This stops all services in the project, including free-tier ones, and resources might be irretrievably deleted. Because notifications are delayed, it does not guarantee you stay under budget.
- Set max instances on every service. Cloud Run now defaults to 100 instances per revision, still far more than a test needs. Lower it with
gcloud run services update SERVICE --max-instances N. The limit can be exceeded briefly during spikes. For Cloud Functions for Firebase, setmaxInstancesin the function options. HTTP requests beyond the limit wait 30 seconds and then get a 429, according to the docs. - Guard against recursion. Firebase warns that a function that writes to the document that triggered it can loop forever, and its own example exits early when nothing has changed. For services that call themselves, pass a depth counter, track visited inputs, and consider task queue functions to throttle the rate.
- Test against the emulator first. Firebase lists fan-out workloads and infinite loops among the common causes of surprise bills and recommends the Local Emulator Suite, which incurs no charges.
What would have caught it sooner
Every warning in this incident arrived late: billing data took about a day, and the Firebase dashboard took longer. In a later edit, Chauhan said Cloud Monitoring alerts reached him with a lag of about 3 to 4 minutes, and that kind of lag is what matters when Firestore is serving a billion reads a minute. CostHex reads usage every minute and alerts via Slack, Discord, Telegram or email with a link to the resource, and it is read-only by default. It currently supports Cloudflare Workers only. Firebase and Google Cloud support is next, but it is not available yet.
Sources
- Founder's post-mortem, Part 1 (archived copy; original blog is offline), Milkie Way (Sudeep Chauhan)
- Founder's post-mortem, Part 2 (archived copy), Milkie Way (Sudeep Chauhan)
- News coverage of the incident, The Register
- Hacker News discussion thread, Hacker News
- Create, edit, or delete budgets and budget alerts, Google Cloud documentation
- Manage spend cap budgets, Google Cloud documentation
- Set up spend caps for Firebase services, Firebase documentation
- Disable billing usage with notifications, Google Cloud documentation
- Set maximum instances for services, Google Cloud documentation
- Manage functions (maximum instances), Firebase documentation
- Cloud Firestore triggers, Firebase documentation
- Avoid surprise bills, Firebase documentation