Skip to content
Service Line · Cloud & AI FinOps

We caught a A$300-a-day runaway
on our own cloud bill.

In July 2026, one usage-billed AI API quietly became ~90% of our own production store’s Google Cloud bill — bots, not customers. One billing-export query found it; one day fixed it. The AI Spend Guardrail Audit installs that exact stack on your estate: SKU-level visibility, budgets that page a human, a daily anomaly scan, and pre-launch gates on every usage-billed API. Fixed fee, quoted on scope in a 20-minute call.

Our own incident, our own numbers
Read-only IAM access
Fixed fee · quote-based
GCP-first, and we say so

The incident this offer is built from.

Every figure in this column is from our own July 2026 production incident — our store, our billing export, our fix. No client’s numbers, no composite anecdote. We sell the guardrail stack because we needed it first.

The full anatomy — the detection query, the four-layer bot-gate, the self-check — is published as a gated report: read the teardown.

A$300/day

One API, burning flat, around the clock

Google’s Retail Search API on our own production ecommerce store — a perfectly flat daily burn with no launch and no traffic spike behind it. Flat was the tell: humans sleep, crawlers don’t.

~90%

Of the entire monthly bill

One usage-billed API had quietly become almost the whole invoice — a ~A$9,000/month run-rate on a store whose human traffic justified a fraction of it. Nothing was broken; every system did exactly what it was configured to do.

~971,762

Search calls, mostly not human

A bot-driven faceted-navigation crawl trap: every filter combination a unique URL, every URL a paid API call. On a public surface, you don’t control who invokes the meter.

1 day

From named SKU to fixed

One BigQuery billing-export query — cost by service and SKU, day over day — named the culprit. A four-layer bot-gate ended it the same day: facet canonicalization, robots.txt, a backend gate serving crawlers a free fallback, and caching. Humans kept AI search; spend fell to noise.

What the audit installs

Five guardrails. Installed, not recommended.

Usage-billed AI APIs invert classic cloud waste: the old failure was paying for idle capacity; the new one is cost that scales with invocations on a surface where you don’t control who invokes. The FinOps Foundation’s State of FinOps 2026 found 98% of practitioners now manage AI spend — this stack is how a team without a FinOps function does it.

01

SKU-level visibility

Billing export to BigQuery, plus the one query that names your top cost SKUs — cost by service and SKU, day over day. The same query that caught our own runaway.

02

Budgets that page a human

Budget alerts at 50, 90 and 100 percent, plus forecast alerts that fire before a threshold is breached — routed to a named person who acts on them, not a dead inbox.

03

Daily anomaly scan

A day-over-day scan by service and SKU across the estate. A runaway stops being a month-end surprise and becomes a same-day incident.

04

Pre-launch gates

Every usage-billed API gets quotas, bot-gating and caching designed before go-live — because AI agents that retry on failure raise the stakes on every public endpoint.

05

Hygiene sweep

Scale-to-zero, right-sizing, storage lifecycle rules, registry cleanup — the quick wins are actioned during the audit itself, not left as a recommendations slide.

Numbers 1–3 caught our runaway. Number 4 would have prevented it. Number 5 pays for the time the audit takes. That order isn’t marketing — it’s the incident timeline.

The honest part: who shouldn’t buy this.

The audit is built for teams running Vertex AI, Retail, Maps, or any per-call API on Google Cloud, spending roughly $5k–50k+ a month, without a dedicated FinOps function. Outside that shape, we’d rather disqualify you here than on an invoice.

Emerge is listed in the Google Cloud partner directory; the work runs on your estate, under your IAM, in your billing export.

Usage-billed AI APIs Vertex AI, Retail, Maps, per-call anything
The core fit

These APIs invert classic cloud waste: the old failure mode was paying for idle capacity; here, cost scales with invocations — and on a public surface you don’t control who invokes. This is exactly the exposure the guardrail stack is built for.

$5k–50k+/mo GCP spend No dedicated FinOps function
Built for you

Big enough for a runaway to hurt, not big enough to staff a FinOps team. The audit installs the function as tooling plus habits — and hands it over rather than renting it back to you.

Dedicated FinOps team Guardrails likely already in place
Run the self-check first

The teardown ends with a 10-point self-check. Score seven or more confident yeses and you don’t need us — we’d rather say that here than discover it two weeks into an engagement.

Multi-cloud estate AWS or Azure carries the spend
GCP-first, honestly

The installed stack is Google Cloud-native: billing export, BigQuery, budget and forecast alerts. The methodology ports; the tooling doesn’t. If your spend lives elsewhere, we say so on the scoping call rather than improvise.

The engagement

One audit. One optional layer on top.

The audit is the product: the guardrail stack installed and handed over in about two weeks. AI Spend Watch exists for teams that want the daily scan read by a human every day — and it is a report subscription, stated plainly, not an on-call service.

The core engagement

Guardrail Audit

The stack, installed and handed over.

A fixed-fee, roughly two-week engagement on your Google Cloud estate: name where the money goes SKU by SKU, install the five guardrails, action the quick wins, and walk your team through running it.

Shape
Fixed fee, ~2 weeks, one Google Cloud billing account
Access
Read-only IAM — billing viewer plus BigQuery read on the export dataset only
Installs
SKU visibility, paging budgets, daily anomaly scan, pre-launch gates, hygiene sweep
Output
Findings doc with named SKUs, the installed stack, a prioritized fix list with quick wins already actioned, handover walkthrough

Best for: Teams on usage-billed APIs who want the runaway class closed once, with the capability left in-house.

Quoted on scope Scope it

AI Spend Watch

The daily scan, with a human reading it.

An optional monthly layer on top of the installed stack: we review the daily anomaly scan, send a monthly trend-and-forecast report, and flag anomalies to a shared channel. A report subscription — not an on-call SLA, and we won’t dress it up as one.

Shape
Monthly subscription, on top of an installed guardrail stack
Cadence
Daily scan reviewed by a human; anomalies flagged to a shared channel
Report
Monthly spend trend and forecast, by service and SKU
Not included
On-call response. We flag; your team acts. That boundary is the product being honest

Best for: Teams that want the scan read by someone whose job it is — without pretending it’s a pager rotation.

Quoted monthly Scope it

There is no rate card on this page because the honest number depends on your project count, spend level, and how many usage-billed APIs you run. Both shapes are quoted on scope in a 20-minute call — and the call can end with “run the self-check, you’re fine.”

How it runs

Read-only in. Capability out.

The usual objection to a cost audit is access, so that’s settled first and in writing: we read your billing data, and we never touch your workloads. Two weeks later the guardrails are installed, the quick wins are done, and the whole thing is yours.

01

Access

Read-only, scoped, in writing

Billing viewer plus BigQuery read on the billing-export dataset — nothing else. No write access to workloads, no deploy permissions, no service-account keys. The access posture is agreed in writing before anyone looks at a number.

02

Audit

Two weeks on the estate

Week one names where the money goes, SKU by SKU, and flags anything that looks like our July incident. Week two installs the five guardrails and actions the hygiene quick wins directly — findings you can verify in your own billing export.

03

Handover

Yours to run

A findings doc with named SKUs, the installed guardrail stack, a prioritized fix list with the quick wins already done, and a walkthrough so your team operates all of it without us. AI Spend Watch is optional after that — never required.

The full teardown

The A$2,000 AI-Spend Runaway — the report behind this page.

The complete incident write-up, with our own production numbers throughout: the anatomy of a paid-API crawl trap, the one BigQuery detection query, the four-layer fix shipped in a day, the guardrail stack that stays on — and the 10-point self-check we now run on every AI workload. If you only take one thing from this page, take the self-check.

The questions a careful buyer actually asks.

What access do you actually need?

Read-only IAM: a billing viewer role plus BigQuery read on the billing-export dataset — and only that dataset. No write access to workloads, no deploy permissions, no production credentials. The scope is agreed in writing before the engagement starts, and you revoke it the day we hand over.

How long does it take?

About two weeks on a single Google Cloud billing account. Week one establishes SKU-level visibility and names where the money goes; week two installs the guardrails and actions the quick wins. It ends with a handover walkthrough, not a dependency.

What does it cost?

The audit is a fixed fee, quoted on scope in a 20-minute call — project count, spend level, and how many usage-billed APIs you run. There is no rate card on this page because the honest number depends on those three things. AI Spend Watch is quoted separately as a monthly subscription.

What if our estate is already clean?

Then the audit’s first finding is that it’s clean, and we say so. The teardown ships with the 10-point self-check we run on our own workloads — score seven or more confident yeses and you don’t need us at all. Run it before you book the call; it’s free and it might end the conversation.

Is AI Spend Watch an on-call service?

No, and we won’t sell it as one. It is a report subscription: the daily anomaly scan reviewed by a human, a monthly trend-and-forecast report, and flagged anomalies posted to a shared channel. Acting on a flag is your team’s call and your team’s hands — the audit will already have given them the runbook.

We’re on AWS or Azure — does this apply?

The methodology ports: SKU-level visibility, budgets that page a human, a daily anomaly scan, pre-launch gates on usage-billed APIs. The installed tooling is Google Cloud-native — billing export to BigQuery, budget and forecast alerts — so this offer is GCP-first. If most of your spend lives elsewhere, we say so on the scoping call instead of stretching the fit.

Twenty minutes to scope it.

Bring your project list and a rough monthly spend. We’ll tell you what the audit would cover, quote it on scope — or tell you to run the self-check and keep your money. Our own runaway ran for weeks before one query ended it; the point of the call is that yours never gets the chance.

Our own incident, our own numbers
Read-only IAM access
Fixed fee · quote-based