Hacker News
Training mRNA Language Models Across 25 Species for $165 — Open-Source Bio-LLM Hits 138 Points
OpenMed published a Hugging Face blog post and open-source release of mRNA language models trained across 25 species using approximately 115 million sequences for just $165 in compute. The NUWA model uses BERT-like architecture with curriculum masked language modeling, covering 19,676 bacterial, 4,688 eukaryotic, and 702 archaeal species. All models, training code, and multi-species datasets released under Apache 2.0/MIT. The extreme cost efficiency ($165 for a production-quality biology foundation model) illustrates how domain-specific AI is becoming accessible to individual researchers.
↳ Follow the thread