An open 10B‑parameter language model, trained from scratch.
Tessera is an independent effort to pretrain Tessera-10B, an open ~10‑billion-parameter language model on Google Cloud TPUs, and to publish the weights and a full technical write-up so others can learn from and build on it.
Mission
Most capable language models are built by large labs, and the details of how they were trained are rarely shared in full. Tessera aims to show, openly and reproducibly, what a single independent developer can do with modern accelerators.
Open by default
Weights, configuration and training recipe are planned for public release, not just a demo.
Reproducible
Documented data processing, hyperparameters and training logs, so that results can be checked and repeated.
Honest reporting
What worked and what didn't, including failed runs and the real compute used.
Technical plan
The current plan. Everything here is planned and may change as small-scale experiments come in.
Model & training
A standard decoder-only transformer written in JAX, with parameters and optimizer state sharded across a 32-chip v6e slice (32 GB HBM per chip), bf16 compute and regular checkpointing to Cloud Storage.
Data (planned)
A mix of openly available, permissively licensed sources: filtered web text, code, math and science writing, and public-domain books. Planned processing: deduplication, quality filtering, removal of personal information, and respecting opt-outs.
Evaluation (planned)
Held-out loss tracked during training, plus standard open benchmark suites run with public evaluation tooling after training. Results will be published as measured.
Compute estimate (back-of-envelope, not a result): ~6 × 10B × 150B ≈ 9×10²¹ FLOPs. At v6e's 918 bf16 TFLOPs per chip peak and an assumed 30–40% utilization, that is roughly 9–12 days on 32 chips.
Roadmap
Phases in order. No phase beyond the first has started.
- Now
0 · Planning
Architecture, data and compute plan; this site; applying for compute.
- Next
1 · Data pipeline & tokenizer
Assemble, deduplicate and filter the corpus; train the tokenizer; document every source.
- Planned
2 · Small-scale experiments
Train small proxy models on a few chips to validate the code, scaling behaviour and hyperparameters before committing to the full run.
- Planned
3 · Main pretraining run
The ~10B-parameter, ~150B-token run on a 32-chip Cloud TPU v6e slice.
- Planned
4 · Evaluation & release
Benchmark the model, then publish the weights, model card and technical write-up.
Openness commitment
The goal is to put something useful back into the open ecosystem.
Weights
Planned release of final (and selected intermediate) checkpoints on Hugging Face under an open license.
Write-up
A detailed technical report: data, architecture, training curves, costs, and lessons learned.
Code
Training and data-processing code planned for release as open source.
About me
Eli Ziv
I'm Eli Ziv, an independent developer working on large language models. I've fine-tuned LLMs on cloud GPUs, built my own pipelines for cleaning and formatting training data, and researched the hardware and memory needs of training at scale. Tessera-10B is my first model trained from scratch, and I'm documenting the whole process in the open.
Contact
Questions, collaboration ideas, or compute partnerships: hello@algebralearningmathonine.onl