Starshine Labs is a research lab studying the frontier.
What a frontier model can actually do is an empirical question, and answering it well is a research problem in its own right. We build evaluations that hold up — and publish the ones everyone should see.
Research
Measurement is a research problem. We study where evals break — contamination, saturation, models that learn the test — and build methods that hold up.
Open benchmarks
We publish benchmarks with full methodology and maintain them as models improve. Results anyone can check are how measurement earns trust.
Partnerships
We work directly with frontier labs, building evaluations for the questions that matter before a model ships.
Thinking about the same problems? We'd like to hear from you — hello@starshinelabs.ai.