AGENTANYWHERE VAANI · INDIC VOICE

Private preview · first pilot in build

Vaani

वाणी

The caller speaks. The recording stays home.

AgentAnywhere Vaani is our Indic voice stack. A voice note or a call comes in — Hindi, Hinglish, Urdu, Tamil, Telugu, Marathi or English — and is understood, masked, classified and answered on Indian infrastructure, on ordinary CPUs. Nothing is kept but an anonymised tally. In private preview; the first pilot deployment is being built now.

Languages
Seven to start — English, Hinglish, Hindi, Urdu, Tamil, Telugu and Marathi.
Runs
Server-side on Indian infrastructure, CPU at runtime. No GPU, no foreign cloud.
Retains
Nothing but an anonymised tally — no audio, no transcript, no identity.
Status
Private preview. Accuracy is being measured on a frozen evaluation set and is not yet published.

In India, the sovereign problem is spoken.

Most of India's conversations with its institutions are not typed. They are said — into a helpline, a bank's WhatsApp number, a clinic's phone, a field officer's handset — in a mix of languages that changes mid-sentence. A recording of that conversation is the most sensitive artefact the institution holds: a voice, a name, an account number, a fear.

The default way to make that recording useful is to send it to a foreign transcription API. That is a data-export event, whatever the contract says. AgentAnywhere Vaani is built so the audio never leaves Indian infrastructure, and so the pipeline keeps nothing it does not need — the transcript is masked before any model reasons over it, and the only thing left afterwards is a count.

Animated diagram of the AgentAnywhere Vaani pipeline: a Hindi voice note enters on the left, its language is confirmed, speech is recognised on Indian infrastructure by an open-licensed external recogniser labelled as interim, Veil irreversibly masks names and numbers, a Tatva model classifies category and urgency with constrained decoding, a templated reply and a pre-rendered voice note go back to the caller, low-confidence or urgent notes go to a human caseworker, and only an anonymised tally is kept.
FIG.30The AgentAnywhere Vaani pipeline — a voice note in, a routed answer out, an anonymised tally left behind. Recognition is an interim external component and is labelled as such.

What happens to every voice note.

Six steps, in the order they run. Each is either ours, or labelled as what it is.

Language

The caller or the worker chooses the language; a check confirms it. Hinglish is treated as its own case, never as Hindi with English words in it.

Recognition — interim

Speech becomes text on our infrastructure. Today this uses an open-licensed external recogniser as an interim component. Our own recogniser replaces it, language by language, when it beats the interim on held-out audio — a rule, not a date.

Masking — Veil

Names, numbers and identifiers are irreversibly masked before any model reasons over the text. The classifier never sees who called.

Classify and route — Tatva

A Tatva model assigns category, urgency and confidence with constrained decoding: a label from your taxonomy, never free text. Low confidence, urgency or any safety flag goes to a person.

Reply

A reply approved by you — text plus a pre-rendered voice note in the caller's language. The copy is customer-owned; the synthesis is an external open-licensed component, labelled as such.

Tally

What remains is an anonymised tally: counters only, with small numbers suppressed. Sessions hold no audio and no transcript.

What is ours, and what is not — said on the page.

The sovereign claim covers the parts that are actually ours. The running system reports the same thing in its own descriptors, so the page and the software cannot disagree.

Ours

The masking (Veil), the classification and routing (a Tatva model trained from scratch), the taxonomy, the reply templates, the safety logic, the anonymised tally, the deployment on Indian infrastructure, and the frozen evaluation set every number will be measured on.

External, interim, labelled

Speech recognition and speech synthesis currently use open-licensed external models. They are described that way everywhere, including inside the running system. Each is replaced by our own model on a published switchover rule — when ours beats the interim on the held-out audio for that language.

Until a language has been measured on the frozen set, no accuracy figure for it appears anywhere. Hinglish is measured explicitly, never inferred from the Hindi number.

Where AgentAnywhere Vaani goes to work.

The stack is general — a voice note in, a routed answer out — and built to be reused. These are the deployments it is being shaped for.

Helplines and CSR programmes

The first deployment: a WhatsApp voice-note helpline where callers describe a problem in their own language, are answered with approved copy, and are handed to a caseworker when it matters. In build now.

PSU banks and banking correspondents

Voice-note requests from customers and correspondents in the field — masked, categorised and routed to the right desk without a transcript ever being stored.

Contact centres and BPOs

Voice notes and IVR fall-through classified before an agent picks up, with nothing retained that a client's data-residency clause would object to.

Insurers

Claims intimation by voice in the language the claimant actually speaks, with health and financial identifiers masked before triage.

Agriculture and rural services

Advisory requests from a farmer's own phone, in their own language — the intake side of what Tatva Kisan is being built to answer.

Health programmes

Appointment and symptom triage lines where the recording must never leave the programme's infrastructure and identifiers must never reach a model.

What we are not claiming yet.

We do not publish an accuracy figure for any language until it has been measured on the frozen evaluation set, and we do not infer Hinglish from Hindi. We do not claim the speech recognition is ours — it is not yet, and the page says so. We do not run the model on the caller's phone; AgentAnywhere Vaani is server-side, on Indian infrastructure. And we do not claim a certification of your deployment: ShepHertz operates a control environment credentialed for SOC 2 and ISO 27001 and independently assessed for HIPAA and GDPR, and our alignment to the DPDP Act is designed-to-support, there to inform your own assessment.

Where AgentAnywhere Vaani is today.

AgentAnywhere Vaani is in private preview. The pipeline runs end to end with fixtures and with the interim components; the first pilot deployment — a voice-note helpline — is being built, and the evaluation set it will be measured on is frozen before the pipeline is tuned, so the number cannot be built to flatter what is already running.

We are taking a small number of design partners with a real voice workload in an Indian language. If your callers speak rather than type, and the recording cannot leave the country, talk to us.

FAQ

Frequently asked questions.

What is AgentAnywhere Vaani?
AgentAnywhere Vaani is ShepHertz's Indic voice stack: a server-side pipeline that takes a voice note or call in an Indian language, recognises it on Indian infrastructure, masks personal data with Veil, classifies and routes it with a Tatva model, answers with approved text and a pre-rendered voice note, and retains nothing but an anonymised tally. It is in private preview; the first pilot deployment is being built.
Which languages does AgentAnywhere Vaani support?
Seven to start: English, Hinglish, Hindi, Urdu, Tamil, Telugu and Marathi. Hinglish is treated and measured as its own language rather than inferred from Hindi. Accuracy per language is being measured on a frozen evaluation set and is not yet published.
Does the audio leave India, and what is retained?
No. Every step runs on Indian infrastructure on ordinary CPUs; no audio is sent to a foreign API. Sessions hold no audio and no transcript, and personal data is irreversibly masked before any model reads the text. What remains afterwards is an anonymised tally of counters only.
Is the speech recognition ShepHertz's own model?
Not yet, and the page says so. Speech recognition and speech synthesis currently use open-licensed external models as interim components, running on our infrastructure. Our own recogniser replaces the interim language by language when it beats it on held-out audio. The masking, classification, routing, taxonomy, templates, tally and evaluation set are ours.
Can AgentAnywhere Vaani work with our WhatsApp number or contact centre?
Yes. The first deployment is a WhatsApp Business voice-note helpline, and the same pipeline accepts voice notes and IVR fall-through from a contact centre. The reply copy and the category taxonomy are yours; low-confidence, urgent or safety-flagged notes are routed to a human you designate.

Bring us a voice workload that cannot leave the country.

A helpline, a WhatsApp number, a claims line — in the language your callers actually speak. We will run it through AgentAnywhere Vaani on Indian infrastructure and show you what is kept: a tally, and nothing else.