Ai2Labs
BenchMIRT: What are LLM benchmarks actually measuring?

BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and…
Read at Ai2 ↗Related

LabsHow a Georgia Tech team used the open Olmo stack to trace social reasoning Ai2

LabsTutorMoments: Do AI tutors know when to help and when to hold back? Ai2

LabsDiScoFormer: One transformer for density and score, across distributions Ai2

LabsWhich tokens does a hybrid model predict better? Ai2

Labsolmo-eval: An evaluation workbench for the model development loop Ai2
