AgentAnywhere Swaraj
How-to15 min read

Train a tiny story model from scratch, and prove what it read

Nine old books of fables and folk tales, one open-source data tool, one short training script and a CPU. About twenty minutes of machine time later you have a small model that writes clumsy story-shaped text, and a manifest that says exactly which stories it was trained on.

AgentAnywhere Research

Pipeline diagram in five columns. Nine public-domain story collections, 672 stories, pass a registry gate that requires source, licence, date acquired, data class and language; nine are admitted and none refused, and 31 stories are held out. The Shuddhi pipeline finds no duplicates and no personal data, drops 35 stories for contamination and 49 through a content screen, and keeps 557 of 641. A build manifest records three hashes. A small story model with 2.6 million parameters and a 4,096-entry tokenizer is trained from random weights on 254,968 tokens.
FIG.65The whole tutorial in one picture, with the real counts from our run: 672 stories prepared, 31 held out, 641 into the pipeline, 557 kept and sealed, one small model trained on them.

Who this is for, and what you need

This is for a developer who has used a language model but never trained one, and who would like to see the whole path once at a size that fits in their head: choosing texts, cleaning them, training a tokenizer and a model, and being able to say afterwards what the model was trained on.

You need Python 3.10 or newer, about 1 GB of disk, and a network connection for the downloads. No GPU and no accounts. On Ubuntu, install the venv module first (sudo apt install python3-venv); on a stock Ubuntu 24.04 machine the first command below fails without it.

Everything here was run for real on 7 October 2026, on rented cloud machines with 8 virtual CPUs and 32 GB of memory. The numbers, hashes and samples below come from a run on a fresh machine that followed only these instructions, and that run's logs, manifests, run card and samples are published in the example's evidence folder.

Step 1: get the example and install the tools

The example is eleven small files (a corpus script, a registry, a filter plugin, a model, a training script, a generation script and their notes) plus a folder of evidence from our run. Shuddhi is installed from its public repository, pinned to the v1.2.0 tag so that your manifest and ours come from the same engine 1.

bash
curl -LO https://agentanywhere.ai/examples/story-model/story-model.tar.gz
tar xzf story-model.tar.gz && cd story-model

python3 -m venv .venv && source .venv/bin/activate
pip install "shuddhi[lid,tokens,extract] @ git+https://github.com/agentanywhere/shuddhi@v1.2.0"
pip install torch --index-url https://download.pytorch.org/whl/cpu
shuddhi doctor

doctor should end with: READY — the pipeline can run in this environment.

Step 2: choose texts you can account for

A model is made of its training data, so the first job is not code. It is deciding which texts you are entitled to use and writing down why.

We set one rule and applied it book by book. A text goes in only if its Project Gutenberg record says it is public domain in the United States, and every named author or translator of the part we use died in or before 1955. That is what we checked. It is not a claim about every country: Project Gutenberg itself tells readers outside the United States to check their own law 6, and so do we.

Nine books passed. They are listed below and, with links, under Sources as 7 to 15. They are nineteenth and early twentieth century retellings: two English versions of Aesop, three collections drawn from the Jataka tales and Indian beast fables, and four of Indian folk tales.

The nine sources, by Project Gutenberg ebook number, with the dates of whoever wrote the text we use:

the nine books
#21     Three Hundred Aesop's Fables, tr. G. F. Townsend (1814-1900), 1867
#28     The Fables of Aesop, retold by Joseph Jacobs (1854-1916), 1894
#36039  The Giant Crab and Other Tales, W. H. D. Rouse (1863-1950), 1897
#30635  The Talking Thrush, W. Crooke (1848-1923) and W. H. D. Rouse, 1899
#57380  Eastern Stories and Legends, Marie L. Shedlock (1854-1935), 1920
#6145   Tales of the Punjab, Flora Annie Steel (1847-1929), 1894
#11310  Hindu Tales from the Sanskrit, S. M. Mitra (1856-1925)
        and Nancy Bell (1844-1933), 1919
#11167  Deccan Nursery Tales, C. A. Kincaid (1870-1954), 1914
#38488  Folk-Tales of Bengal, Lal Behari Day (1826-1894), 1883

Text only. No illustrations, prefaces by other hands, notes or indexes.

What we left out, and why that matters

The rule cost us five good candidates, and the reasons are the useful part.

A well-known illustrated Aesop for children from 1919 has unsigned text. It is very probably fine. But "very probably fine" is an argument, and when a text's status needs an argument we leave it out. Two further collections, one of Aesop and one of Jataka tales, have a translator or author whose year of death is missing from the catalogue record, so the rule could not be applied to them.

The most instructive case was a famous 1892 anthology of Indian fairy tales. Its title page carries one editor, who died in 1916. Its notes at the back show that the tales were taken from about a dozen other collections and translators, several of whose dates we could not confirm from catalogue records. An anthology is not one text; it is as many texts as it has sources. We left it out, and a second anthology with it. In one book we kept, we cut the introduction, which was written by someone who died in 1961.

A provenance tool cannot make these decisions for you. It can make sure somebody made them, and wrote them down.

Step 3: prepare the corpus

Shuddhi reads plain text with one blank line between documents, one file per source. The script downloads the nine files, removes the Gutenberg header and licence block (the Gutenberg licence asks for that 5), cuts each book down to its stories, and writes one story per document. Every twentieth story is held out and never written to the training files.

bash
python prepare_corpus.py

aesop_townsend   pg21      313 stories   298 to corpus   15 held out
aesop_jacobs     pg28       82 stories    78 to corpus    4 held out
giant_crab       pg36039    28 stories    27 to corpus    1 held out
talking_thrush   pg30635    43 stories    41 to corpus    2 held out
eastern_stories  pg57380    29 stories    28 to corpus    1 held out
punjab_tales     pg6145     43 stories    41 to corpus    2 held out
hindu_tales      pg11310    92 stories    88 to corpus    4 held out
deccan_nursery   pg11167    20 stories    19 to corpus    1 held out
bengal_tales     pg38488    22 stories    21 to corpus    1 held out

672 stories: 641 for training in corpus/, 31 held out in heldout/eval-set.jsonl.

Step 4: the registry and the provenance gate

Shuddhi will not read a file until someone has declared where it came from. You can watch that happen. Ask it to scaffold a registry from the folder, then check the scaffold:

bash
shuddhi init --corpus corpus --out scaffold.json
shuddhi check --registry scaffold.json

registry: corpus  sha256=2a9935a6459ac2e7…
accepted: 0
refused: 9
  ✗ aesop_jacobs: untagged shard: missing provenance field(s) ['source', 'license',
    'date_acquired', 'data_class', 'language'] — refusal is the default
  ...

Exit code 2. An empty field is a refusal, so the scaffold cannot become a corpus by accident.

The example ships a registry.json with the fields filled in from the checks in step 2. One entry looks like this, and the other eight follow the same pattern:

json
{
  "shard_id": "aesop_townsend",
  "path": "corpus/aesop_townsend.txt",
  "source": "Project Gutenberg ebook #21: Three Hundred Aesop's Fables, translated by George Fyler Townsend (1867). https://www.gutenberg.org/ebooks/21",
  "license": "Public domain in the USA (Project Gutenberg record); Gutenberg header and licence text removed",
  "date_acquired": "2026-10-07",
  "data_class": "public",
  "language": "eng"
}

shuddhi check --registry registry.json → accepted: 9, refused: 0.

Step 5: screen the stories, with a filter you can read

These are old books, and old children's books are not always gentle. We searched the story text for about a hundred terms and read the passages around the strongest hits. That turned up no slurs, but it did turn up real cruelty in places: a burial alive, corpses, a flaying, a plan to take one's own life. Most of it sits in the two longest folk-tale collections.

Shuddhi's built-in lexicon screen is tuned for overtly abusive text and only fires when a document contains at least two distinct listed terms at a minimum density. We wanted something stricter and simpler, so we used the tool's plugin interface 4: a filter of about thirty lines that drops any story containing a term from a short list. The list covers four categories (graphic violence and cruelty, self-harm, adult themes, and dated language about caste, faith or illness) and has 34 entries. We do not reproduce it here; it is a plain text file in the example, and you should read it and change it.

Two things are worth knowing about this screen. It is blunt: it was written from that search of the books, it removes the cruellest stories, and it does not make the rest gentle. A fable in which the wolf eats the lamb stays in. And it is on the record: a plugin must declare everything that changes its verdicts, and that declaration (here, a hash of the term list) is folded into the build's configuration hash. Change one term and the receipt changes.

python
class StoryScreen:
    name = "story-screen"
    version = "1.0.0"

    def __init__(self):
        path = os.environ.get("STORY_SCREEN_TERMS", TERMS_FILE)
        with open(path, encoding="utf-8") as f:
            rows = (line.strip().lower() for line in f)
            self.terms = sorted({r for r in rows if r and not r.startswith("#")})
        self.rx = re.compile(r"(?<!\w)(" + "|".join(map(re.escape, self.terms)) + r")(?!\w)", re.I)

    def identity(self):
        # Everything that changes a verdict: the exact term list.
        sha = hashlib.sha256("\n".join(self.terms).encode("utf-8")).hexdigest()
        return {"terms_sha256": sha, "n_terms": len(self.terms)}

    def check(self, text):
        m = self.rx.search(text)
        return f"screened term: {m.group(1).lower()}" if m else None

story_screen/shuddhi_story_screen.py. Install it with: pip install -e story_screen

Step 6: build the corpus with one command

The pipeline command runs every stage in the right order: measure each file, mint the corpus hash, cluster near-duplicates, apply the filters, write the cleaned text, the manifest and a report. We give it the held-out stories as an evaluation set, name the plugin, and ask for a log of every dropped document.

bash
pip install -e story_screen
shuddhi pipeline --registry registry.json --out shuddhi-out/ \
    --eval-set heldout/eval-set.jsonl --plugin story-screen \
    --no-perplexity --log-drops

corpus_build_hash: 24c7066b0df2f16ad60568bdb41ae6426dc51084fe3973e89ce6335d35da5a7f
docs: 641 total / 641 unique (global exact-dup 0.00%)
near-dup: 0 clusters, 0 docs to drop (largest cluster 0), 0.0 min
  filter plugin: story-screen 1.0.0
done   aesop_jacobs: kept 75, dropped {'plugin:story-screen': 3}
done   aesop_townsend: kept 286, dropped {'plugin:story-screen': 12}
done   bengal_tales: kept 2, dropped {'contamination': 14, 'plugin:story-screen': 5}
done   deccan_nursery: kept 1, dropped {'contamination': 18}
done   eastern_stories: kept 27, dropped {'plugin:story-screen': 1}
done   giant_crab: kept 23, dropped {'plugin:story-screen': 4}
done   hindu_tales: kept 82, dropped {'contamination': 1, 'plugin:story-screen': 5}
done   punjab_tales: kept 27, dropped {'contamination': 1, 'plugin:story-screen': 13}
done   talking_thrush: kept 34, dropped {'contamination': 1, 'plugin:story-screen': 6}
filtered_build_hash: f95a59f41d3849de18cb265b0cdcfd91fc4e16b1a21048928e2b153dde564acd
kept 557 docs; dropped {'exact_dup': 0, 'near_dup': 0, 'quality': 0, 'perplexity': 0,
  'toxicity': 0, 'contamination': 35, 'pii': 0, 'plugin:story-screen': 49}

Output trimmed to the lines that matter. Wall-clock time on 8 CPU cores: 1.8 seconds.

What each stage did to this corpus

Provenance gate: nine files declared, nine admitted, none refused.

Duplicates: none. 641 stories, 641 unique, and no near-duplicate clusters. Two translations of the same Aesop fable are different sentences, and the near-duplicate stage is built to catch copies and templates, not paraphrase 3.

Quality and toxicity screens: nothing dropped. These catch broken text and overt abuse, and a corpus prepared from printed books has neither.

Personal data: zero redactions. The detectors look for email addresses, phone numbers, card numbers and identity numbers, and a book printed in 1894 contains none. A zero here is a measurement, not a skipped step.

Content screen: 49 stories dropped by the plugin, counted on their own line in the manifest.

Contamination: 35 stories dropped, and this is the result worth a closer look.

The contamination check found something we had not planned for

We held out 31 stories so that we could measure the model on text it had never seen. Shuddhi's contamination stage enforces that: any training story that shares a run of eight consecutive words with a held-out story is dropped. We expected it to catch nothing.

It removed 18 of the 19 remaining Deccan Nursery Tales and 14 of the Folk-Tales of Bengal. The drop log gives the reason for each drop but not the matching passage, so the example includes a short script that finds it:

bash
python why_dropped.py

 18 x deccan_nursery  shares "once upon a time there was a town" with heldout:deccan_nursery:20
  6 x bengal_tales    shares "my story endeth the natiya thorn withereth etc" with heldout:bengal_tales:20
  5 x bengal_tales    shares "here my story endeth the natiya thorn withereth" with heldout:bengal_tales:20
  2 x bengal_tales    shares "sons and daughters here my story endeth the" with heldout:bengal_tales:20
  1 x bengal_tales    shares "begetting sons and daughters here my story endeth" with heldout:bengal_tales:20
  1 x hindu_tales     shares "and it was a long time before he" with heldout:hindu_tales:60
  1 x punjab_tales    shares "you don t suppose i am going to" with heldout:punjab_tales:20
  1 x talking_thrush  shares "let me go let me go but the" with heldout:talking_thrush:20

Why we let those 35 drops stand

Every Deccan tale opens with the same sentence about the same town, and every Bengal tale closes with the same rhyme. The held-out story from each book carries the formula too. A model trained on eighteen copies of an opening line would score well on the nineteenth for a reason that has nothing to do with learning to tell a story. That is exactly what a contamination check is for, and it noticed something about our own evaluation design that we had not.

The last three drops are different: three ordinary turns of phrase, eight words long, that happen to occur in two stories. An eight-word rule is strict, and it does not know a formula from a coincidence.

We could have removed the formulas from the text, or chosen different stories to hold out, until the number looked better. We kept the rule and lost the stories, because the point of holding stories out is a number you can trust. It cost about a fifth of the corpus by size, and most of two books. You might decide otherwise for your own model. Either way the manifest records what was done.

One stage we switched off, and a limitation in v1.2.0

Shuddhi also has a perplexity filter, which scores each document against a small character model of your own corpus and drops the strangest one percent. On a corpus this small the tool itself says not to use it: with the filter on, v1.2.0 warns "Skip the perplexity filter for a corpus this size (--no-perplexity)" for any file under 50 documents or 200 KB, which is all nine of ours, and its quickstart gives the same advice for anything under a few thousand documents 2.

We ran it once anyway to see what would happen. With several small files in one language, v1.2.0 took its cutoff from the one file large enough to supply one and applied it to all nine, and ordinary stories were dropped as a result. That is a limitation of v1.2.0 with multiple small files in one language, and a fix is planned. The build in this guide therefore runs with --no-perplexity, as the tool's own guidance says, and the manifest's configuration records no perplexity cutoff.

Step 7: read the manifest

The build writes shuddhi-out/build/BUILD-MANIFEST.json. Three values in it, plus one hash per output file, are the receipt. If you followed the steps with the same nine source files, yours are identical to these, character for character:

receipt
corpus_build_hash     24c7066b0df2f16ad60568bdb41ae6426dc51084fe3973e89ce6335d35da5a7f
filter_config_sha256  34b083f52b964c23519becbcb2e80788ce91c514563d7755faf0faa9f12601a0
filtered_build_hash   f95a59f41d3849de18cb265b0cdcfd91fc4e16b1a21048928e2b153dde564acd

kept_docs             557
dropped_by_reason     contamination 35 · plugin:story-screen 49 · all others 0
pii_redactions        0

From shuddhi-out/build/BUILD-MANIFEST.json, run of 7 October 2026.

What each hash means

corpus_build_hash answers what was measured. It is a hash over the set of all 641 unique stories that passed the gate, before any filter. It does not depend on file order or on how the work was split up.

filter_config_sha256 answers how the selection was made. It covers every threshold and switch, the near-duplicate list, the lexicon, and each plugin's declared identity, so our 34-term list is in there.

filtered_build_hash answers which stories were kept. It is the same kind of hash as the first, taken over the 557 stories that survived, and it covers those documents and nothing else. The configuration is not mixed into it. That is deliberate: anyone holding the same source files can recompute it without our settings, and two builds that arrive at the same stories by different routes share this hash and are told apart by the configuration hash beside it.

One line in the v1.2.0 manifest is worded wrongly, and you will see it in your own file: the hash_definition field says "Chain: parent hash + filter config sha -> this hash". Nothing is folded in. The meaning is the one above: which documents were selected, with the configuration recorded separately in filter_config_sha256.

The manifest also records a sha256 for each cleaned text file. Those tell you what bytes were written, which matters when a filter rewrites text rather than dropping it. Alongside the manifest, the pipeline writes REPORT.md, a draft training-content summary laid out against the European Union's template for general-purpose models, with the nine sources in a table and the gaps it cannot fill marked as gaps. Our toy model needs no such filing; it is there so you can see what one looks like.

Step 8: the model

The model is a small decoder-only transformer: four layers, four attention heads, a width of 192 and a context of 256 tokens, which comes to 2,615,424 parameters. Nothing is loaded from anywhere. Every weight starts as a small random number.

python
CONTEXT, LAYERS, HEADS, WIDTH, DROPOUT = 256, 4, 4, 192, 0.2


class Block(nn.Module):
    def __init__(self):
        super().__init__()
        self.ln1, self.ln2 = nn.LayerNorm(WIDTH), nn.LayerNorm(WIDTH)
        self.attn = nn.MultiheadAttention(WIDTH, HEADS, dropout=DROPOUT, batch_first=True)
        self.mlp = nn.Sequential(nn.Linear(WIDTH, 4 * WIDTH), nn.GELU(),
                                 nn.Linear(4 * WIDTH, WIDTH), nn.Dropout(DROPOUT))

    def forward(self, x, mask):
        h = self.ln1(x)
        x = x + self.attn(h, h, h, attn_mask=mask, need_weights=False)[0]
        return x + self.mlp(self.ln2(x))


class StoryModel(nn.Module):
    def __init__(self, vocab):
        super().__init__()
        self.tok_emb = nn.Embedding(vocab, WIDTH)
        self.pos_emb = nn.Embedding(CONTEXT, WIDTH)
        self.drop = nn.Dropout(DROPOUT)
        self.blocks = nn.ModuleList(Block() for _ in range(LAYERS))
        self.ln = nn.LayerNorm(WIDTH)
        self.head = nn.Linear(WIDTH, vocab, bias=False)
        self.head.weight = self.tok_emb.weight          # tied weights
        self.apply(self._init)

    def forward(self, idx, targets=None):
        n = idx.shape[1]
        mask = torch.triu(torch.ones(n, n, dtype=torch.bool, device=idx.device), 1)
        x = self.drop(self.tok_emb(idx) + self.pos_emb(torch.arange(n, device=idx.device)))
        for block in self.blocks:
            x = block(x, mask)
        logits = self.head(self.ln(x))
        if targets is None:
            return logits
        return F.cross_entropy(logits.view(-1, logits.size(-1)), targets.view(-1))

model.py, lightly trimmed (the weight-initialisation helper is omitted here).

Step 9: train, starting from the receipt

The training script does three things before it trains anything. It reads the manifest. It hashes each cleaned file and refuses to start if the bytes on disk are not the bytes the manifest recorded. Then it trains a byte-pair tokenizer of 4,096 entries on those same stories and nothing else.

python
manifest = json.load(open(os.path.join(args.build, "BUILD-MANIFEST.json")))
stories = []
for shard, entry in sorted(manifest["per_shard"].items()):
    path = os.path.join(args.build, f"{shard}.filtered.txt")
    data = open(path, "rb").read()
    recorded = entry["emitted_file"]["sha256"]
    if hashlib.sha256(data).hexdigest() != recorded:
        raise SystemExit(f"{path} does not match the manifest. Refusing to train.")
    stories += [s for s in data.decode("utf-8").split("\n\n") if s.strip()]

From train.py. A corpus edited after the build cannot be trained on by accident.

Then the loop: 2,000 steps of 32 windows of 256 tokens, with the loss on the held-out stories measured every 100 steps. The script keeps the checkpoint with the best held-out loss.

bash
python train.py

corpus: 557 stories, filtered_build_hash f95a59f41d3849de18cb265b0cdcfd91fc4e16b1a21048928e2b153dde564acd
held out: 28 stories the model never trains on
tokens: 254,968 train, 17,285 held out, vocabulary 4096
model: 2,615,424 parameters, trained from scratch on cpu
step   100  train 6.877  held-out 6.013      63s  saved
step   200  train 5.411  held-out 5.272     122s  saved
step   500  train 4.435  held-out 4.761     301s  saved
step  1000  train 3.740  held-out 4.548     603s  saved
step  1300  train 3.506  held-out 4.532     781s  saved
step  1500  train 3.388  held-out 4.541     902s
step  1900  train 3.296  held-out 4.530    1143s  saved
step  2000  train 3.297  held-out 4.531    1204s
done in 1204.3s; best held-out loss 4.530 at step 1900; wrote model/

Log trimmed to nine of twenty lines. 28 held-out stories, not 31: three contained screened terms.

Training loss, held-out loss, and what the gap means

Two numbers matter, and they should always be read together. The training loss is how surprised the model is by the stories it is learning from. The held-out loss is how surprised it is by 28 stories it has never seen, stories that Shuddhi has checked share no eight-word run with the training set.

For the first thousand steps both fall. Then the held-out loss flattens at about 4.53 while the training loss keeps falling, to 3.30 by the end. From that point the model is not getting better at stories. It is getting better at these stories: it is memorising.

That is the expected result when 2.6 million parameters meet 254,968 training tokens. There are about ten parameters for every token of text, which is plenty of room to store the text instead of learning from it. A held-out loss of 4.53 means that on unseen text the model is, on average, about as uncertain as if it were choosing among ninety-odd equally likely next tokens. That is a model that has learned the shape of a fable, not its sense.

If you only ever look at the training loss, a model that is memorising looks like a model that is improving.

The script also writes model/run.json, a run card. It records the architecture, that the weights started from scratch, the full loss log, and the three hashes and nine file hashes of the corpus. The model now carries the answer to "what was this trained on?" with it.

Time and cost

On the 8-core machine, everything before training (installing the tools, nine downloads, the corpus script and the Shuddhi build) took about two minutes, of which the build itself was 1.8 seconds. Training took 1,204 seconds, which is 20 minutes 4 seconds. Generating six samples took six seconds.

From the moment the machine was created to the moment it was destroyed was 22 minutes, and the provider charged ₹6.64 for it. A laptop with a recent processor will do the same work; expect it to be slower with fewer cores.

If you do rent a machine, destroy it when you finish and check the provider's listing to see that it is gone. A machine shut down from the inside may still be billed. We learned that on an earlier attempt, and it cost rather more than the run did.

Step 10: generate stories, and check them for copying

The generation script prompts the model with a title and prints what it writes. For every sample it also reports the longest run of consecutive words that appears word for word in the training text. A short run is ordinary: any English writer reuses five-word phrases. A long run means the model is reciting.

bash
python generate.py
python generate.py --prompt "Once upon a time there was a crow"

Sampling settings: temperature 0.65, top 40 tokens, up to 220 tokens per sample.

Three real samples

Six samples were generated with the default prompts, and every one was read and screened before any was printed. None contained anything we would call unsuitable; one that used the word "slave", which the corpus contains because Aesop's own story is that of one, was left out. Here are three, exactly as generated, with line breaks removed. They were chosen to show the range, not the best case.

One of the better ones. Prompt: "The Tortoise and the Peacock". A KID, having been a little clearing at the Tortoise, he looked up and saw him, and said: “I have been sitting on the sand, and I am sure you had to eat.” Next day the Tortoise had a Tortoise, and he had fluttered down, and went to the Tortoise. “I have been bitten by your wings, you here,” said the Tortoise, “and you put me on the sand and then I shall have been a good meal to-pots.” “You shall have come to eat,” said the Tortoise. “I would have been a turning by and I am a good friend. I shall have to come to-morrow, and I will do it again.” The Tortoise was so delighted, and he said: “Sir I am sure to get him.” “See, and then, Dan I can help you?” “No, friend,” said the Tortoise, “we can come along with me.” “Well, then,” said the Tortoise, “and I have made you

164 words. Longest run copied from the training text: 5 words ("on the sand and I"). This is invention, such as it is: the rhythm of a fable, dialogue that opens and closes its quotation marks, a tortoise that talks to itself, and no story.

Mostly recall. Prompt: "The Clever Sparrow". A FLE was prowed with a tree, and when he saw a fil, said, 'I am indeed, but the Sparrow was hungry, and when a feast began to eat me, as you will you, and when you can eat me, or a khichrî eat; and if you had to eat me, and it so you can wash beak and feet, and eat you the khichrî as a khichrî eat the Cow, and said-- 'Your name is milk, my name is Grass, For I give me a nice khichrî eat; Though I really don't you can mean, For you can't know What the Sparrow can mean, For I'm sure, For I'm sure, Ap, I'm sure, For I'm remarkably clean!' So the Sparrow sat, and said, 'Certainly I'm remarkably clean, and I'm clean, and I'm remarkably clean!' But the Sparrow paid the Cow, and said, 'Certainly you must go to the Deer, for the Cow, and I will give you a nice khichrî eat

174 words. Longest copied run: 11 words ("…t know what the Sparrow can mean, for I'm sure"). We label this one as recall, not invention. The word "Sparrow" pulled the model into a rhyme from one Punjab tale, "The Sparrow and the Crow", and it is reassembling that rhyme from memory, khichrî and all. A repeated verse in a small corpus is the first thing a small model memorises.

A weak one. Prompt: "The Fox and the Crow". A PHERTERTTARTERTTTTTERTTTTTTTTTTTERST was a WING, having a certain sand, and on a Trees’s por and often bought a wood. But the Crow came up and a house, and called that his son as a Crow. When the Crow was the Crow, she had not enough to give her some fer to do. The Crow’s pots, and made a pure of pentereaks penprine and mars, but it was so much money to be enough to eat. The Crow said, "Why do you do you do you give me a grain of corn? Do you can do not eat yourself to-tree." The Crow did not know what the Swan, but he was fond of corn, and said, "Monkey, you must be very fond of corn." The Monkey agreed to tell this; but the Crow was not like a nice of corn, but he used to the

152 words. Longest copied run: 6 words ("the Crow did not know what"). "The Fox and the Crow" is a real title in the training text, twice, so we expected recitation here. We got the opposite. The likely cause: one of the two Aesop translations opens every fable with a few words in capitals, capitals are rare enough in this corpus that the tokenizer breaks them into fragments, and the model stutters before it finds its feet. It never mentions a fox.

What this model can and cannot do

What it learned

The surface of a fable: a title, an animal with a capital letter, a short exchange of dialogue, a turn.

Which words keep company: crows with corn, tortoises with sand and water.

Punctuation habits, including which quotation marks each source book uses.

What it cannot do

Tell a story. Nothing it writes has a beginning, a problem and an end, and it does not keep track of who is speaking.

Stay on your prompt, answer a question, or follow an instruction. It only continues text.

Be trusted with an audience. The screen removed the cruellest source stories, not every hard edge, and a model this small can stitch harmless words into something odd. Do not put it in front of children unsupervised.

Did it reproduce?

The run described here was itself the reproduction test. We had already run the whole thing once on another machine of the same size and written down its counts, hashes and losses. We then started a fresh machine and followed only the README in the example folder. The one stumble was the one in the prerequisites: the venv module was missing on stock Ubuntu.

The corpus reproduced exactly. All nine source files had the same sha256, the three manifest hashes matched character for character, and so did the sha256 of every cleaned file. That is the property the manifest exists for: two people who start from the same sources and the same configuration can prove they trained on the same corpus without sending each other the corpus.

The training reproduced as well. Everything in the record we had committed before the second run matched exactly: the source hashes, the three manifest hashes, the nine cleaned-file hashes, the counts, the losses at the six logged steps written down there, and the best held-out loss. The remaining logged steps and the six samples also matched what we had recorded from the first run in our working notes and in the draft of this post.

Only the clock differed: the two machines were a couple of minutes apart. On a different processor, thread count or library version, expect small differences in the loss and different samples.

The manifests, the run card, the full training log and all six samples from this run are in the example's evidence folder. The manifest records the name of the machine that built it; in the published copy that field is redacted, and the home directory in two logs with it. If you publish a manifest of your own, look at that field first.

Where to go next

Give it more to read. The single biggest improvement is more text. Add another public-domain collection, apply the same rule to it, add one entry to the registry and rebuild. The hashes will change, as they should.

Change the screen. Edit the terms file, rebuild, and watch filter_config_sha256 change while you compare what was dropped. Then decide whether you agree with our list.

Try the other route on contamination. Hold out different stories, or strip the opening and closing formulas in the corpus script, and see how many of the 35 come back and what happens to the held-out loss.

Shrink the model. With this little text, a smaller model may generalise as well and memorise less. The run card makes the comparison honest: same filtered_build_hash, different architecture.

Read the rest of the documentation. The quickstart and user guide cover every stage and its limits, and the plugin guide covers filters like ours.

The same discipline is what matters at scale. A two-megabyte corpus and a two-hundred-gigabyte one are accounted for the same way: declare the sources, measure everything, record what was removed and why, and cite the hash.

Frequently asked questions

Can I train a language model from scratch without a GPU?

Yes, if the model and the corpus are small. The model in this guide has 2,615,424 parameters and is trained on 254,968 tokens of text. On a rented machine with 8 CPU cores, 2,000 training steps took 20 minutes 4 seconds, and the whole run cost under ₹7. The result is a toy that writes short, clumsy story-like text, which is the honest limit of a model and corpus this size.

What does 'trained from scratch' mean here?

Every weight in the model starts as a small random number and the tokenizer is trained on the same corpus. Nothing is downloaded or fine-tuned from an existing model. The run card written by the training script records this as initialised_from: scratch, alongside the hashes of the corpus.

How do I prove what a model was trained on?

Build the corpus with a tool that records it. Shuddhi writes a build manifest containing a corpus build hash (what was measured), a filter configuration hash (how the selection was made) and a filtered build hash (which documents were kept), plus a sha256 for each output file. The training script in this guide refuses to start unless the files match the manifest and copies the hashes into a run card kept with the model. Anyone with the same source files and configuration can recompute the hashes.

Does the filtered build hash include the filter settings?

No. The filtered build hash covers which documents were selected and nothing else. How they were selected is recorded separately in filter_config_sha256. Two builds that select the same documents by different settings share a filtered build hash and are told apart by the configuration hash. A description line in the v1.2.0 manifest words this wrongly; the behaviour is as described here.

Are the stories used in this guide free to use?

Each of the nine books is recorded by Project Gutenberg as public domain in the United States, and every named author or translator of the text used died in or before 1955. That is what was checked. Copyright law differs by country, so check the law where you are before reusing them. Several other candidate books were left out because their status could not be confirmed to the same standard.

Why did the contamination check remove 35 stories?

Thirty-one stories were held out for evaluation, and any training story sharing a run of eight consecutive words with a held-out story is dropped. Two of the source books begin or end every story with the same formula, so the held-out story from each book matched 18 and 14 of its siblings. Three more matched on an ordinary eight-word phrase. The drops were left in place so that the held-out loss measures text the model has not seen.

Is the story model safe to use with children?

No, not unsupervised. It is a learning exercise. The source books are nineteenth and early twentieth century retellings, a term-list screen removed the cruellest stories but not every hard edge, and a model this small produces text that is often incoherent. Every sample printed in this guide was read and screened before it was printed.

Sources

Shuddhi is open source under Apache-2.0. Documentation links point at the v1.2.0 tag used in this guide. Gutenberg records were read on 7 October 2026; rights were checked as described in step 2, and readers should check the law where they are.

  1. 1ShepHertz — Shuddhi (शुद्धि), source repository.
  2. 2Shuddhi — Quickstart.
  3. 3Shuddhi — User Guide.
  4. 4Shuddhi — Extending Shuddhi: filter plugins.
  5. 5Project Gutenberg — The Project Gutenberg License.
  6. 6Project Gutenberg — Permissions, licensing and other common requests.
  7. 7Aesop, tr. George Fyler Townsend — Three Hundred Aesop's Fables, Project Gutenberg #21.
  8. 8Joseph Jacobs — The Fables of Aesop, #28.
  9. 9W. H. D. Rouse — The Giant Crab, and Other Tales from Old India, #36039.
  10. 10William Crooke and W. H. D. Rouse — The Talking Thrush, and Other Tales from India, #30635.
  11. 11Marie L. Shedlock — Eastern Stories and Legends, #57380.
  12. 12Flora Annie Steel — Tales of the Punjab, #6145.
  13. 13S. M. Mitra and Nancy Bell — Hindu Tales from the Sanskrit, #11310.
  14. 14C. A. Kincaid — Deccan Nursery Tales, #11167.
  15. 15Lal Behari Day — Folk-Tales of Bengal, #38488.
Topicstrain a language model from scratchsmall language model tutorialtraining data provenanceShuddhi tutorialpublic domain training databuild manifest for training datatrain a tokenizer

Written by

AgentAnywhere Research

The team that builds the platform and the models

AgentAnywhere Research writes about the platform, the model families and the trust layer we build and run in India. Where a figure is ours, it says what it covers; where something is a demonstration or in preview, it says so.

All articles →
Announcement

Introducing Tatva Edge: the model that says no

Tatva Edge is an on-device AI model that turns a plain-language instruction into exactly one command a machine understands, or refuses. Offline, in India's languages, with a person in command. Waitlist open.

Siddhartha Chandurkar5 min read