Reasoning Evaluation
Separate reasoning from recall.
THE FRONTIER AI EVALUATION GAUNTLET
Can You Help Evaluate the
Intelligence That Comes Next?
APPLICATION WINDOWSeptember 16, 2026 · 11:59 PM IST
01 / THE OTHER HALF OF THE RACE
Building increasingly capable AI models is only one part of the race. The other part is learning how to measure them.
You are not here to build another wrapper. You are here to separate capability from coincidence, make failures reproducible and discover what intelligence can actually do.
02 / CHALLENGE AREAS
A focused experiment.
A defensible finding.
Separate reasoning from recall.
Test the complete engineering loop.
Follow decisions across long horizons.
Find confident answers without evidence.
Build tests that resist shortcuts.
Probe the boundary of reliable behavior.
Measure across modalities.
Inspect actions, recovery and state.
Change the conditions. Test again.
Surface consequential failure modes.
Design judgments that hold up.
Make experiments repeatable.
Turn anomalies into evidence.
Give every finding a rerunnable path.
Measure what happens after the demo.
03 / YOUR FIRST TEST
Apply. Receive a progression decision within 24 hours. If advanced, your team gets exactly 24 hours from qualifier issuance to submit its evidence.
T + 0
UP TO 24 HOURS
CLOCK STARTS
24 HOURS
The normal maximum application-to-submission process is approximately 48 hours. A qualifier is not guaranteed: the leader may instead receive a non-progression decision. No further edits are guaranteed after submission; late submissions are recorded and flagged.
04 / THE GOA 15
Top 10 announced September 18.
Remaining slots follow extended technical review.
05 / THE PRIZES
Eligible recipients receive six months of access, subject to account eligibility, product and geographic availability, applicable service terms, and the organizer’s provisioning or reimbursement mechanism.
The Codex benefit is funded independently by the HACK THE FIRST AGI 1.0 organizing team. Product references do not imply sponsorship or endorsement by OpenAI or any other company.
06 / THE RESIDENTIAL FINAL
15 teams. 30–45 builders. Approximately 3–4 days inside a focused frontier-AI research challenge.
October 1–15, 2026 window. Exact final dates and venue will be communicated privately. Maintain availability until dates are confirmed.
Subject to published approval and reimbursement rules. Standard funded accommodation is one room per team, with legitimate accessibility, medical, safety and privacy exceptions.
Read the full funding policy ↗15 → 3 → 1
Applications close September 16, 2026 · 11:59 PM IST