Enterprise AI CX. Your data never leaves your project.
A production-grade customer-experience agent your customers can type to, talk to, or send
a photo — grounded on your real systems, never on guesswork. Run it on Google’s frontier
models under enterprise terms, on open weights you own inside your own cloud project, or
fully on-premise where no cloud is permitted.
Private corpus
Permission
Retrieve
Reply
Review
Customer outcome
Trust rings from a private corpus through permissioned retrieval to a customer reply with human review.
Illustrative workflow. Licensing, permissions and feasibility are confirmed during discovery. This is not a claim of a native connector or a delivered client result.
Apache 2.0 weights you own16 GB VRAM class deployNative voice + vision256K context windowIn-country KSA verifiedGCC PDPL alignedDirectory-listed Google Cloud partnerQuote-based engagement
Apache 2.0 weights you own16 GB VRAM class deployNative voice + vision256K context windowIn-country KSA verifiedGCC PDPL alignedDirectory-listed Google Cloud partnerQuote-based engagement
Open weights changed the procurement question.
Until mid-2026, a capable multimodal CX assistant meant sending customer conversations
to a model vendor, token by token. Google’s Gemma 4 release ended that trade-off: a
12-billion-parameter model that hears, sees, and reasons — small enough to run on
hardware your project already has access to, licensed so that you own it.
Every figure on this band is verified against Google’s published Gemma 4 developer
guide and model card (June 2026) — and re-verified at scoping.
Apache 2.0
Weights you own
Gemma 4 12B ships under an open licence. No per-token fee to a model vendor, no obligation to send a single customer utterance outside your boundary.
16 GB VRAM
One enterprise GPU
Google’s own guidance: small enough for 16 GB of VRAM. In practice, a single NVIDIA L4 serves it — which changes what “private deployment” costs.
Voice + vision
Natively multimodal
The encoder-free architecture ingests speech and images directly — the largest audio-capable model in the Gemma 4 family. A customer can talk to it or send a photo.
256K
Context window
Enough working memory for long support conversations, full policy documents, and the retrieved knowledge that grounds every answer.
The sovereignty ladder
One architecture. Three answers to “where does our data live?”
The agent design is identical at every rung — the same tool-grounded loop, the same
channels, the same insights. What changes is where the model runs, and who holds the keys.
Managed CX
Frontier capability, enterprise terms.
Gemini on Vertex AI, running in your own Google Cloud project against regional endpoints.
Model
Gemini on Vertex AI
Boundary
Your GCP project, under Google Cloud’s published Vertex AI data-governance commitments
Residency
Vertex regional endpoints, region-pinned
Voice
Native audio in, Cloud TTS out
Best for: Retail and commerce CX, and any program where capability speed matters more than weight ownership.
Pricing follows the engagement, not a rate card: every tier is scoped and quoted against
your integration surface, channels, and regulatory posture during the Assessment.
The residency map, stated plainly.
Most vendors answer the residency question with a hand-wave. We answer it with region
codes — including the regions where the honest answer is “not there, use the
Sovereign tier.” That precision is the point.
GPU quota is confirmed for your specific project during the Assessment — documented
availability and granted quota are not the same thing, and we treat them accordingly.
me-central2 · DammamSaudi Arabia
In-country inference, verified
L4 and G4 GPU capacity is documented for this region. The Private tier runs Gemma 4 12B on GKE inside the Kingdom.
me-central1 · DohaQatar
No GPU capacity today
Managed tier serves via regional endpoints. A hard in-country inference requirement routes to the Sovereign tier.
no regionUnited Arab Emirates
No Google Cloud region exists
An in-country UAE mandate is answered honestly: the Sovereign tier, on your own infrastructure.
asia-south1 / south2 / southeast1 +Nearest Cloud Run GPU
Mumbai, Delhi, Singapore
Where in-country is preferred but not mandated, the Private tier can run serverless on Cloud Run GPU in six documented regions.
Reference implementation
Live proof, not a slide.
“Ask Aqua” runs today on aquora.ae —
an Emerge-operated UAE commerce property with a catalogue of roughly six thousand
products. It is the Managed tier of this exact architecture, in production: an agentic
tool loop in which every product, price, and order status the agent shows is grounded
in a live system result. It cannot invent a SKU.
The Private and Sovereign tiers swap the brain for Gemma 4 12B.
The proven design does not change.
Live catalogue search and comparison across ~6,000 products
Photo-to-product visual search
Voice in and out, with a dual speech-recognition fallback path
Cross-sell and personalised recommendations
Order-status lookup, gated behind order number and email
PII-masked logging, full transcripts, assisted-conversion attribution
An admin insights dashboard over every conversation
Canned replay — no live model on this page
One interaction, replayed.
The loop below is the whole architecture in ten seconds: the agent reasons, calls tools,
retrieves grounding, and answers from system results. On the Private tier, the identical
loop runs on weights you own, inside your project.
Replay · recorded reference interaction
My pool pump has started rattling — do you carry a replacement impeller for a 1.5 HP
AquaPro, and could I collect it today?
Working…
Understanding the request
catalogue.search
stock.lookup — Dubai warehouse
Impeller compatibility guide
AquaPro 1.5 HP series accepts the genuine impeller kit and one certified compatible alternative; both fit without re-plumbing.
Warehouse stock feed — live
Both kits held at the Dubai warehouse; same-day collection window closes 17:00.
Yes — your 1.5 HP AquaPro takes two impeller kits: the genuine part and a certified compatible alternative. Both are in stock at the Dubai warehouse right now, and same-day collection is open until 5pm. I can reserve one under your order email — would you like the genuine kit or the alternative?
Every card above is a system result, not model memory — the loop that cannot invent a SKU.
The same Crawl–Walk–Run discipline as every Emerge program.
Sovereignty programs fail when the architecture is chosen before the regulation is
read. So the engagement starts with a fixed-scope Assessment — and every phase after
it is gated, named, and quoted before it begins.
01
Crawl
Data-Residency AssessmentTailored quote
A fixed-scope diagnostic. Output: your regulatory posture mapped, a tier recommendation with evidence, an integration inventory, and the target architecture. Built so procurement can sign it quickly.
02
Walk
Build & launchTailored quote
The agent built and grounded on your real systems — knowledge base, order or case APIs, escalation paths. Channels to match your posture: web, WhatsApp where permitted, email. Named KPIs and phase gates throughout.
03
Run
OperateTailored monthly quote
A managed monthly cadence: model updates, prompt and tool tuning, evaluation reviews, and an SLA. Your team keeps custody; ours keeps it sharp.
The questions procurement actually asks.
Who owns the model?
On the Private and Sovereign tiers — you do. Gemma 4 is released under Apache 2.0, so the weights sit in your own registry, and any fine-tune we build on your data is yours outright. There is no licence to renew and no vendor who can switch it off.
Is a 12B open model good enough for customer-facing work?
The honest answer: a frontier managed model is stronger on open-ended tasks. A CX agent, however, is not open-ended — it is tool-grounded and narrowly scoped, and the architecture does the heavy lifting. That is why the Assessment includes an evaluation gate on your actual use-cases before you commit to a tier. We publish no benchmark claims; we test on your traffic.
What about Arabic?
We scope and evaluate Arabic quality per use-case during the Assessment rather than making a blanket claim. If the evaluation gate says the quality is not there for your dialect and domain, we tell you — and design the fallback before launch, not after.
Can we run WhatsApp on the Sovereign tier?
WhatsApp messages transit Meta’s Business Platform by design — for most sovereign deployments that is disqualifying, so we default those programs to web, voice, and kiosk channels. If your regulator signs off on WhatsApp for specific journeys, we wire it. We would rather surface that tension in scoping than discover it in an audit.
Start where the regulation starts.
The Data-Residency Assessment is a fixed-scope Crawl phase: your regulatory posture, the
right tier with evidence, and a procurement-ready architecture — before you commit to a
build.