Independent · Planned open release

An open 10B‑parameter language model, trained from scratch.

Tessera is an independent effort to pretrain Tessera-10B, an open ~10‑billion-parameter language model on Google Cloud TPUs, and to publish the weights and a full technical write-up so others can learn from and build on it.

Mission

Most capable language models are built by large labs, and the details of how they were trained are rarely shared in full. Tessera aims to show, openly and reproducibly, what a single independent developer can do with modern accelerators.

Open by default

Weights, configuration and training recipe are planned for public release, not just a demo.

Reproducible

Documented data processing, hyperparameters and training logs, so that results can be checked and repeated.

Honest reporting

What worked and what didn't, including failed runs and the real compute used.

Technical plan

The current plan. Everything here is planned and may change as small-scale experiments come in.

~10Bparameters, decoder-only transformer
~150Btraining tokens
JAXtraining stack
TPU v6eCloud TPU Trillium
32 chipsplanned main training slice

Model & training

A standard decoder-only transformer written in JAX, with parameters and optimizer state sharded across a 32-chip v6e slice (32 GB HBM per chip), bf16 compute and regular checkpointing to Cloud Storage.

Data (planned)

A mix of openly available, permissively licensed sources: filtered web text, code, math and science writing, and public-domain books. Planned processing: deduplication, quality filtering, removal of personal information, and respecting opt-outs.

Evaluation (planned)

Held-out loss tracked during training, plus standard open benchmark suites run with public evaluation tooling after training. Results will be published as measured.

Compute estimate (back-of-envelope, not a result): ~6 × 10B × 150B ≈ 9×10²¹ FLOPs. At v6e's 918 bf16 TFLOPs per chip peak and an assumed 30–40% utilization, that is roughly 9–12 days on 32 chips.

Roadmap

Phases in order. No phase beyond the first has started.

  1. Now

    0 · Planning

    Architecture, data and compute plan; this site; applying for compute.

  2. Next

    1 · Data pipeline & tokenizer

    Assemble, deduplicate and filter the corpus; train the tokenizer; document every source.

  3. Planned

    2 · Small-scale experiments

    Train small proxy models on a few chips to validate the code, scaling behaviour and hyperparameters before committing to the full run.

  4. Planned

    3 · Main pretraining run

    The ~10B-parameter, ~150B-token run on a 32-chip Cloud TPU v6e slice.

  5. Planned

    4 · Evaluation & release

    Benchmark the model, then publish the weights, model card and technical write-up.

Openness commitment

The goal is to put something useful back into the open ecosystem.

Weights

Planned release of final (and selected intermediate) checkpoints on Hugging Face under an open license.

Write-up

A detailed technical report: data, architecture, training curves, costs, and lessons learned.

Code

Training and data-processing code planned for release as open source.

About me

EZ

Eli Ziv

I'm Eli Ziv, an independent developer working on large language models. I've fine-tuned LLMs on cloud GPUs, built my own pipelines for cleaning and formatting training data, and researched the hardware and memory needs of training at scale. Tessera-10B is my first model trained from scratch, and I'm documenting the whole process in the open.

Contact

Questions, collaboration ideas, or compute partnerships: hello@algebralearningmathonine.onl