Kodawire — SEE EVERY SIDE. KNOW THE TRUTH.Kodawire — SEE EVERY SIDE. KNOW THE TRUTH.Home
Log in
HomeWorldPoliticsBusinessTechHealthScienceClimateEducation
Kodawire — SEE EVERY SIDE. KNOW THE TRUTH.Kodawire — SEE EVERY SIDE. KNOW THE TRUTH.

Balanced coverage of the week's biggest story. See every side. Know the truth.

By signing up for the Story of the Week newsletter you agree to our privacy policy.

Product

  • For You
  • Trending
  • Blindspot
  • Local
  • Compare
  • Brief
  • Trust

Topics

  • World
  • Politics
  • Business
  • Tech
  • Health
  • Science

Regions

  • Africa
  • Europe
  • Asia-Pacific
  • Americas
  • Middle East

Company

  • About
  • Contact
  • All topics
  • My Bias
  • Privacy
  • Terms & DMCA

KODAWIRE

KODAWIRE

© 2026 Kodawire

AboutContactPrivacyTermsTop
/
/
Back

Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

1 sources

Google Developers Blog (XX)2h ago

Center

Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

The MaxText team successfully reproduced AI2’s OLMo 3 7B language model from scratch on Google Cloud TPUs using JAX/XLA, precisely matching the original PyTorch-on-GPU reference across pre-training and mid-training stages on all held-out evaluations. The implementation achieved up to 57.4% Model Flops Utilization (MFU) and demonstrated robust infrastructure portability by surviving mid-run cluster resizes and cross-generation TPU shifts without requiring recipe alterations. Crucially, the exercise proved the necessity of comprehensive held-out validation by catching a silent data-loader memorization bug that artificially depressed training loss and would have otherwise faked a performance win.

Original

Spectrum

L 0%C 100%R 0%