TATVA NANO · TRAINED FROM SCRATCH IN INDIA
Private preview · capability proof · larger checkpoints in trainingTatva Nano
नैनोToken one was ours.
Tatva Nano is the smallest member of the Tatva family and the clearest proof of its lineage: a compact language model trained from zero in India — no foreign base, no inherited weights — on a corpus we cleaned and sealed ourselves. It writes India's languages in their own scripts, the Dravidian languages included, and runs fully offline on a laptop CPU. In private preview as a capability proof; larger checkpoints are in training.
- Lineage
- From scratch. No base model, no inherited weights; every training document's licence recorded.
- Languages
- Writes Indian languages in their own scripts — Devanagari, Dravidian, Bengali.
- Runs
- Fully offline on a laptop-class CPU. No network, no GPU, no cloud.
- Status
- Capability proof in private preview. Evaluations under way; no accuracy figure published yet.
Why a small model matters more than a big claim.
Almost every model sold as sovereign inherits its weights from somewhere. That is a reasonable engineering choice and an awkward regulatory one, because the first question a defence programme or a bank's model-risk team asks — *what did it learn from?* — usually has to be answered with “we do not fully know”.
Tatva Nano exists to make the other answer possible. It was initialised at random and trained on a corpus we assembled, cleaned and sealed ourselves, with the licence of every document recorded and a build receipt every training run cites. We can show the corpus, the manifest, the tokenizer and the run. It is small because a proof should be legible, not because small is the ambition.
What Tatva Nano is.
A from-scratch lineage
Random initialisation, our corpus, our run. Nothing inherited, so there is nothing to explain away in front of an auditor.
A tokenizer built for Indian scripts
Devanagari, the Dravidian scripts and Bengali are first-class in the vocabulary, not an afterthought bolted onto an English tokenizer.
A sealed corpus with receipts
The training data passed through the Shuddhi data factory: cleaned, filtered under a hashed configuration, and sealed with a build receipt the run cites.
Offline, on a laptop
The checkpoint runs on a laptop-class CPU with the network off. The demonstration is given that way on purpose.
What it does today, and what it does not.
A capability proof is only useful if its edges are drawn honestly.
Tatva Nano does
Continue a sentence in fluent script in a major Indian language — Hindi, Marathi, Bengali, Tamil, Telugu, Kannada, Malayalam among them — from a few words of prompt.
Run fully offline on ordinary hardware, and carry a complete provenance record from corpus to checkpoint.
Serve as the checkpoint the rest of the from-scratch family grows from.
Tatva Nano does not
Act as an assistant or a knowledge base. It will invent facts and it will ramble; it is a language model, not an answer engine.
Carry a published accuracy figure yet. Evaluations are under way on a frozen set; the number appears when it is measured, not before.
Stand in for the larger checkpoints. Usefulness beyond the proof comes with the next rungs, which are in training on the same lineage.
Who asks for it, and why.
The buyers who care about Tatva Nano are the ones who get asked where a model came from.
Defence and strategic programmes
A model whose composition can be shown rather than asserted — and that runs air-gapped, because it was never dependent on anything outside the box.
PSU banks and model-risk teams
Model risk management wants training-data provenance on record. With Tatva Nano the record exists end to end, and Model Hub keeps it.
Government and public-sector programmes
Indic-first by construction, trained in India on Indian-language text, with no foreign base to disclose or to depend on.
Indian-language products
The lineage Tatva Kisan and the AgentAnywhere Vaani classifier draw on: the same tokenizer, the same corpus discipline, the same receipts.
See it run.
The demonstration does one honest thing. You give Tatva Nano the start of a sentence in any Indian language — a few words in Hindi, Tamil or Telugu — and it continues it in fluent script, on a laptop with the network switched off. It is a five-minute proof that the lineage is real, and we give it in person or on a call.
Where Tatva Nano is today.
Tatva Nano is in private preview as a capability proof. The from-scratch training programme continues on the same lineage — larger checkpoints are in training and will be described here when they are measured, not before. If provenance is the question your reviewers ask first, this is the model to see.
Related
Where Tatva Nano fits.
FAQ
Frequently asked questions.
- What is Tatva Nano?
- Tatva Nano is the smallest member of AgentAnywhere's Tatva family: a compact language model trained from zero in India — no base model, no inherited weights — on an Indian-language corpus cleaned and sealed by the Shuddhi data factory. It writes India's languages in their own scripts and runs fully offline on a laptop CPU. It is in private preview as a capability proof.
- Is Tatva Nano a fine-tune of another model?
- No. Tatva Nano was initialised at random and trained on ShepHertz's own corpus with an Indic tokenizer built for Indian scripts. Every training document's licence is recorded and each training run cites a build receipt, so the lineage can be shown end to end rather than asserted.
- Which languages can Tatva Nano write?
- Major Indian languages in their own scripts, including Hindi, Marathi and Bengali in Devanagari and Bengali script, and all four major Dravidian languages — Tamil, Telugu, Kannada and Malayalam. Given the start of a sentence, it continues it in fluent script.
- Can Tatva Nano answer questions or act as an assistant?
- Not today, and the page says so. Tatva Nano is a capability proof of a from-scratch lineage: it continues text fluently but will invent facts. Usefulness beyond the proof comes with the larger checkpoints now in training on the same lineage.
- Does Tatva Nano have published accuracy numbers?
- Not yet. Evaluations are under way on a frozen evaluation set, and no figure is published until it has been measured there. What is available now is the provenance record — corpus, manifest, tokenizer and training run — and a live offline demonstration.
Ask us where the model came from. Then watch it run offline.
Tatva Nano is in private preview as a capability proof. If provenance is the first question your reviewers ask, we will show you the corpus, the receipts and the checkpoint — on a laptop with the network off.