Marin's 535 billion parameter model training run launched with support of The Jen-Hsun and Lori Huang Foundation GPU gift

Open Athena is thrilled to announce that, with a generous gift of compute from The Jen-Hsun and Lori Huang Foundation, we have kicked off Marin's largest training run yet—a 535 billion total parameter large language model (LLM)—and the largest open and live training run in history.

Science progresses through exploration and observation, hypothesis and experiment, debate and consensus. It is the development of science in the open that enables the brightest minds everywhere to participate in this dialogue and by doing so improve it. The science of artificial intelligence is no different, and Marin exists so that this science can both benefit from the public and benefit the public.

Some scientific discoveries—the microscope, the transistor, genome sequencing—have themselves become tools in better answering scientific questions. Without AI that can be understood, inspected, and perturbed, both the science of artificial intelligence and AI for science will take a back seat to private research and development. Marin exists to ensure this vital tool can be used by the public for academic research, and that academia can participate in and contribute to the effort.

"The frontier of AI shouldn't be built inside a black box. Marin's 'open lab' approach, opening up every code pipeline, experiment, and training log, is crucial for the AI ecosystem," said David Siegel, Founder and Chairman of Open Athena. "We are deeply grateful for the Huang Foundation's partnership in advancing a more transparent and accessible future for scientific research."

Our model in training is a 535B total parameter, 23B active parameter mixture of experts model, with 18 trillion tokens of data, including agentic, coding, and scientific data. We have publicly predicted what the performance of our model will be after its pretraining completes on December 1st, as measured by a specific score on a standard benchmark (Paloma). The prediction is based on runs 300x smaller than this one: we will continue to improve the model as it runs, so this will serve as a lower bound on performance.

Scaling is crucial: It is only by building and studying state-of-the-art large language models that scientists can fully understand the technology and best benefit from its use. However, reaching this AI frontier requires significant computational power, which puts it out of reach of most academic labs, nonprofit groups, and even most commercial labs. Thanks to the generous grant from the Huang Foundation, the Marin project will be able to study and train models at a scale that approaches the frontier, empowering researchers to transparently study the hardest and most interesting problems in AI.

"AI is becoming one of the most important instruments of scientific discovery. For researchers to advance the science of AI – and use AI to advance every field of science – they need access to frontier-scale models they can build, study, and understand. The compute the Foundation is providing to Marin puts that capability into the open, giving researchers everywhere the opportunity to participate in shaping the next generation of AI," said Jensen Huang.

Beyond language modeling, we have two non-language research projects taking place in the open, on the Marin platform: MarinDNA (GitHub) and MarinFold (GitHub). MarinDNA, for example, produced a 1B genomic language model competitive with Evo 2 40B, while using ~1,980× fewer training FLOPs and scoring variants ~2,330× faster. We are exploring collaborations in other areas including climate, plant genomics, and materials science that will also be developed on the Marin platform.

Marin began in 2024 at Stanford University, created by David Hall and Percy Liang. Beyond the explicit and process knowledge generated through discovery and practice in the open, Marin provides model weights, training checkpoints, data recipe, and code for infrastructure, data pipelines, training, and inference, and convenes the discussion and community. Marin scaled dense LLMs from 8B parameters to 32B parameters, running on donated compute from Google's TPU Research Cloud, a scaling suite "Delphi", and most recently released a 67B-A2B MoE model on 10T tokens, trained on v4 TPUs. Open Athena is a non-profit founded in 2024 by David Siegel, Jeff Hammerbacher, Mike Abbott, and Jared Crooks to bring the capabilities of the AI frontier to academia. Today, all engineers working full-time on Marin are employed by Open Athena. If you would like to work on Marin at Open Athena, send a note to info@openathena.ai. To learn more about how Marin works, we suggest that you start with mumwelt, a CLI and set of skills to ask questions of the Marin corpus.

Cite this post

@misc{athena2026_huang_foundation_marin_535b_training_run,
  author = {Athena, Open},
  title = {Marin's 535 billion parameter model training run launched with support of The Jen-Hsun and Lori Huang Foundation GPU gift},
  year = {2026},
  month = {sep},
  howpublished = {\url{https://www.openathena.ai/blog/huang-foundation-marin-535b-training-run/}},
  note = {Open Athena Blog}
}