Ai2Labs
olmo-eval: An evaluation workbench for the model development loop

olmo-eval is an open evaluation workbench that helps model developers add, run, and analyze benchmarks across changing LLM checkpoints, extending OLMES from final-score reproducibility into the…
Read at Ai2 ↗Related

VoicesOpen-world evaluations for measuring frontier AI capabilities AI Snake Oil

VoicesUnderstanding the 4 Main Approaches to LLM Evaluation (From Scratch) Ahead of AI

VoicesUsing LLM-as-a-Judge For Evaluation: A Complete Guide Hamel Husain

LabsIntroducing AIMIP: The AI weather and climate model intercomparison project Ai2

VoicesLLM Research Papers: The 2026 List (January to May) Ahead of AI
