Carbon-3B: A 3B DNA Foundation Model That Matches Evo2-7B at 150x the Speed
Carbon-3B matches Evo2-7B on sequence recovery, variant-effect prediction, and motif-perturbation discrimination while generating DNA over 150 times faster.
8 articles tagged with #transformers.
Carbon-3B matches Evo2-7B on sequence recovery, variant-effect prediction, and motif-perturbation discrimination while generating DNA over 150 times faster.
How TTT-E2E achieves constant inference latency regardless of context length by treating long context as a learning problem rather than an architecture problem.
How the same transformer architecture powering GPT learned to predict molecular properties by treating chemistry as a language problem
How researchers adapted BERT for molecular property prediction, turning SMILES strings into drug discovery insights
Pedro Domingos proposes that neural networks and symbolic AI are the same mathematical operation - a logical rule can be equivalently written as a tensor equation in Einstein summation notation. If true, we've been building separate tools for problems that share identical structure.
GPT-4's 128K context window? It only uses about 10% effectively. Google's TITANS architecture introduces test-time memory learning that outperforms GPT-4 on long-context tasks with 70x fewer parameters.
Nested Learning: The Illusion of Deep Learning Architectures - A comprehensive guide to the arXiv paper revealing how neural networks learn at multiple timescales through hierarchical optimization.
Traditional readability formulas miss the mark. Modern embedding models can capture semantic nuance and syntactic structure, but do they actually predict complexity better?