SQA Agenthon
Agenthon Submissions Deadline: October 12th, 23:59 AoE
(October 13th at 7:59 AM New York, October 13th at 12:59 PM London, October 13th at 7:59 PM Singapore)
Announcing Competition Awards!

NeurIPS 2026 Competition Track

Agenthon 2026

Verifiable AI for quantitative finance.

A four-track competition testing whether AI agents can produce finance outputs that survive automated, leakage-controlled, cheat-resistant evaluation.

Agenthon 2026 logo
Submission Docker agent
Admissibility g0-g3 gates
Ranking Track metrics

Awards

$1,500, $1,000 and $500 in every track.

The first, second and third teams in each of the four tracks win cash awards: $12,000 in total.

We're delighted to welcome NVIDIA, Bloomberg and AllianceBernstein as sponsors of Agenthon 2026. Thank you for your support, for helping us organize the competition, and for collaborating with us on its science and task development. Meet our sponsors

Key dates

Key dates.

Registration runs from 17 August to 12 October 2026. Registration and Development close together at 23:59 Anywhere on Earth (AoE, UTC−12) on 12 October. The last Development runs start by 20:00 UTC on 12 October: an upload that has not started by then is not run. The evaluation fleet is in scheduled maintenance on 13 October from 08:00 to 12:00 UTC, when the Final + Verification phase opens.

  1. 17 Aug – 12 Oct

    Registration

    Sign up and form your team.

  2. 28 Aug – 12 Oct

    Development

    Public practice repositories, iteration, and the validation leaderboard.

  3. 30 Sep

    Call for papers closed

    Papers for the NeurIPS workshop were due at 23:59 AoE.

  4. 13 – 25 Oct

    Final + Verification

    One submission per entered track is evaluated on sealed held-out units; organizers rerun leading submissions and review reproducibility within the same phase.

  5. Sat 12 Dec

    Workshop at NeurIPS, Atlanta

    The Agenthon workshop, with final presentations by the winners, in Atlanta, Georgia.

The competition closes with a workshop and final presentations from the winning teams at NeurIPS in Atlanta, Georgia, on Saturday 12 December. All dates are given in the official rules, which govern if anything here is out of step.

Teams

Enter as a team of one to three.

You can compete alone or with up to two teammates. You join one team and stay on it, and that one team can enter as many of the four tracks as it wants.

Registration is per team, and only registered teams can submit solutions or appear on the leaderboard. Your team uses one account on the competition platform, chosen by the team — the FAQ explains what that means for everyone else on the team.

Have a question about entering? Start with the FAQ.

Core question

Can AI agents produce finance answers that can be checked by machine?

Agenthon extends the Alphathon program into its first NeurIPS edition. The competition keeps the four-track structure and hard finance setting, while adding sealed held-out data, automated leakage controls, reproducible reruns, and public leaderboards.

Agents in all four tracks run in sandboxed Docker containers without general internet access. Supported Coding, Forecasting and Explainability submissions can call House Nemotron through a restricted connection. Simulation has no network access. Build required dependencies and permitted artifacts into your image before evaluation.

Four tracks

Same competition spine, four finance problems.

T1 Coding

Quant-finance coding agents

Build a Docker agent that solves quantitative finance coding tasks under pytest and financial-invariant checks.

Verb
solve
Metric
pass@1
Gate
pytest + invariants
T2 Forecasting

Reasoning-augmented time series

Forecast future panels using time-series data plus a time-stamped text corpus. The score evaluates the forecast distribution; information uplift is a research aim.

Verb
forecast
Metric
CRPS composite
Gate
as-of cutoff + calibration
T3 Simulation

Accelerated market simulation

Submit an ABIDES-compatible simulator that is faster while preserving matching-engine semantics and market stylized facts. Development results are practice feedback, not controlled Final timing.

Verb
simulate
Metric
events/sec
Gate
semantic regression
T4 Explainability

Evidence-grounded prediction

Predict labels, values, or rankings over tabular entities, supported by a frozen evidence corpus. Follow the track guide for prediction, evidence and reasoning evaluation.

Verb
analyze
Metric
Track 4 composite
Gate
faithfulness + embargo

Protocol

Submit, check, score.

Each unit is checked for admissibility and evaluated using its track's scoring rules. Participant failures remain in the evaluation denominator; organizer faults are handled separately. If two Final submissions finish a track with the same ranking score, the tie is broken in favour of the one uploaded earlier.

01

Submit

Upload the toolkit-generated ZIP identifying your agent image and proving team registration.

02

Check

g0-g3 verify integrity, schema, cutoff/resource rules, and domain semantics.

03

Score

Track-specific scoring produces Development feedback. Track 1 uses pass@1 without a confidence interval.

Public/private firewall

Transparent grading, sealed answers.

Each track has a public practice repo and a private sealed exam repo. Public repos include practice tasks, examples, local checks, available baselines and participant docs. Private repos hold evaluation tasks, reference answers, scoring material and audit records.

Public practice Private exam
Public practice units Private-test held-out units
Examples, local checks and available baselines Oracle solutions and final scorer
Manifest and canary safety checks Canary registry and audit logs

Leaderboard

Competition scores, one track at a time.

Development leaderboards on agenthon.net are public and visible to everyone. Each track uses its own scale. Higher is better for the displayed leaderboard scores; the raw Forecasting loss is converted for this display. Use CodaBench for uploads and processing status. Follow the logged-in website for submission opening status. From 5 October 2026, 00:00 AoE (12:00 UTC), the Track 1 board ranks only runs uploaded after 25 September 2026 at 05:00 UTC, for every team (Track 1 README, rules 8 and 9).

The leaderboard is updated periodically from the competition platform.

T1 Coding T2 Forecasting T3 Simulation T4 Explainability

T1 Coding

T1 Coding: teams 1–25 of 101

Rank Team Score
1 Kaile 0.7674
2 Saifuddin 0.7209
3 Pluvia 0.5349
4 Proxima 0.5000
5 Arsen Ibragimov 0.4535
6 DeepSucc 0.3837
6 Chirayu 0.3837
8 Yan Su 0.3605
9 RandomGuyHavingFun 0.3488
10 MindAlchemist 0.2907
11 RizYL 0.2791
12 abc123 0.2674
12 Nana Champions 0.2674
14 dbuddha 0.2558
15 Cornfield Chase 0.2326
15 Formaggio 0.2326
17 Autonomous Alpha 0.2209
17 fiftyfifty 0.2209
19 Liqourena 0.2093
19 Tengen Toppa 0.2093
19 Pluto 0.2093
22 Anonymous 0.1977
22 in3lab 0.1977
22 APTAPT 0.1977
22 Quiet Signal 0.1977

T1 Coding: teams 26–50 of 101

Rank Team Score
22 Husky River Trading 0.1977
27 eagle 0.1860
27 Submartingale 0.1860
27 CMCCGDYDDICT 0.1860
27 Make It Make Cents 🙏 0.1860
27 eoo0m 0.1860
32 Just-in-Time 0.1744
32 VerityLovityFalsityCruelty 0.1744
32 BoGo 0.1744
32 Paragon 0.1744
36 Undefined 0.1628
36 TappuKiMKC 0.1628
36 Proof of Alpha 0.1628
36 AbangFrabi 0.1628
36 verifiquant 0.1628
36 FedorAzarov 0.1628
36 Powerhouse 0.1628
36 Eyyyyyy 0.1628
44 Sean 0.1512
44 Four Sigma 0.1512
44 Money Miner 0.1512
44 1991 0.1512
44 Give me a job 0.1512
49 gheware 0.1395
49 Warkop PuteraPrakoso 0.1395

T1 Coding: teams 51–75 of 101

Rank Team Score
49 Masato 0.1395
52 iluss 0.1279
52 Duo_tech 0.1279
52 Jade 0.1279
55 MicroCat 0.1163
55 hnhparitosh 0.1163
55 Agent_007 0.1163
55 always win 0.1163
55 Nietzsche 0.1163
60 Jin & Pei 0.1047
60 TEst 0.1047
62 UNIVERSEPLAYER 0.0930
62 Sir Thaddeus 0.0930
62 HTHY 0.0930
62 Banana Kingdom 0.0930
66 Agent Alpha 0.0814
67 The Unscheduled Penguins 0.0698
68 Apex 0.0581
68 The Mockingjay 0.0581
70 GAOKUO 0.0465
71 attention is qkv 0.0349
72 volatility_labs_stony_brook 0.0233
72 Wealth-Building Lab 0.0233
74 Urithiru 0.0116
— Aamir Abdul Azeez — not scored

T1 Coding: teams 76–100 of 101

Rank Team Score
— AgentReady — not scored
— Orrery — not scored
— DeltaNLP — not scored
— 3226 — not scored
— AlbertQuant — not scored
— Verifiable Capital — not scored
— a-team — not scored
— BetterCallSaul — not scored
— FAWDA — not scored
— 9:00pm — not scored
— C_cake — not scored
— huhudawang — not scored
— Quant Science Flow — not scored
— lingsio — not scored
— phtree — not scored
— DaoyuShu4094 — not scored
— Future Ability — not scored
— RRRRRR — not scored
— Echoing — not scored
— Luca — not scored
— Artifical General Invesment — not scored
— Flowfront — not scored
— nogugu — not scored
— Outsiders — not scored
— team — not scored

T1 Coding: teams 101 of 101

Rank Team Score
— IMAnonymous — not scored

No team matches that search.

T2 Forecasting

T2 Forecasting: teams 1–25 of 124

Rank Team Score
1 DuML -1.2002
2 huhudawang -1.2509
3 Yan Su -1.3403
4 Ywin -1.5483
5 Flowfront -1.5548
6 alpha-extractor -1.5725
7 Paragon -1.5813
8 Made in Heaven -1.6351
9 Samoyed -1.6713
10 Chirayu -1.7656
11 Not Today -1.7697
12 Quiet Signal -1.7917
13 Cornfield Chase -1.7953
14 HBIT -1.8222
15 F1_1 -1.8367
16 Jin & Pei -1.8424
17 Proof of Alpha -1.8515
18 labubu666 -1.8693
19 CMCCGDYDDICT -1.8758
20 Pawsibble -1.9221
21 VerityLovityFalsityCruelty -1.9314
22 Apex -1.9618
23 HackStreet Boys -1.9722
24 Nexus of f(x) -2.0134
25 Warkop PuteraPrakoso -2.0303

T2 Forecasting: teams 26–50 of 124

Rank Team Score
26 mikelou1 -2.0579
27 Just-in-Time -2.0898
28 TheAxiom -2.0940
29 DKYnumber1 -2.1028
30 Autonomous Alpha -2.1097
31 Jigglyduck -2.1111
32 Super Debugging -2.1651
33 Saifuddin -2.1769
34 spidy -2.1999
35 Sapien -2.2129
36 urayaha -2.2155
37 lifeisluck -2.2172
38 TradeFlare -2.2310
39 LastDigitsOfPi -2.2338
40 Masato -2.2448
41 HG Market System -2.2556
42 quack -2.2893
43 Money Miner -2.2898
44 Sitadel Insecurities -2.2962
45 JJ LAB -2.3146
46 Formaggio -2.3310
47 iluss -2.3524
48 abc123 -2.3567
49 5 o'clock -2.3690
50 time -2.3694

T2 Forecasting: teams 51–75 of 124

Rank Team Score
51 fiftyfifty -2.3772
52 pranshu rastogi -2.3774
53 1991 -2.3901
54 AsOf Quant -2.3990
55 Tengen Toppa -2.4107
56 Bayes Street -2.4108
57 Auror -2.4301
58 QIE DOG -2.4329
59 Javier Amo -2.4429
60 Dog parents -2.4467
61 lingsio -2.4480
62 hnhparitosh -2.4484
63 C_cake -2.4513
64 in3lab -2.4570
65 LiLaiLai -2.4637
66 randomwalk -2.4666
67 DoublePai -2.4681
68 Aamir Abdul Azeez -2.4722
69 jusep -2.4751
70 JessH -2.4795
71 ZZKK -2.4795
72 In Sync -2.4848
73 Natpaphon -2.4891
74 Verifiable Capital -2.4918
75 GAOKUO -2.4932

T2 Forecasting: teams 76–100 of 124

Rank Team Score
76 Quant Science Flow -2.4933
77 SMTM -2.4938
78 krnb -2.4961
79 Richard Zhu -2.5069
80 LIZARD -2.5086
81 1111 -2.5090
82 PSS_ -2.5096
83 Garros Solo Lab -2.5139
84 GXR123 -2.5169
85 XIII -2.5228
86 Nietzsche -2.5272
87 Yongchang -2.5273
88 Try hard -2.5347
89 AbangFrabi -2.5381
90 Banana Kingdom -2.5388
91 Orrery -2.5402
92 Jade -2.5458
93 JiaxinLi -2.5482
94 pluuuuuus -2.5521
95 Husky River Trading -2.5567
96 Probably Right -2.5643
97 Liqourena -2.5769
98 Launchlane Forecast -2.5799
99 iKun-XinglinForrest -2.5994
100 YIC Quant -2.6067

T2 Forecasting: teams 101–124 of 124

Rank Team Score
101 Luca -2.6137
102 lwz888 -2.6141
103 Undefined -2.6151
104 ARKANE -2.6357
105 OOBs -2.6373
— Reference forecaster M0 reference, not ranked -2.6412
106 SoloTraveler -2.6765
107 HaHaHa -2.6829
107 ah_a1g -2.6829
107 Give me a job -2.6829
110 OffTheTape -2.7369
111 BigWorld -2.7476
112 katharsis -2.7795
113 K-uant -2.7985
114 PA-Agent -2.8354
115 team -2.9680
— 9:00pm — not scored
— ZOLDIK — not scored
— Urithiru — not scored
— phtree — not scored
— RRRRRR — not scored
— Hatenabase — not scored
— Edge-X — not scored
— Artifical General Invesment — not scored
— Outsiders — not scored

No team matches that search.

T3 Simulation

T3 Simulation: teams 1–25 of 48

Rank Team Score
1 fuzz 231024.2407
2 VerityLovityFalsityCruelty 222058.1706
3 abc123 221088.2954
4 IMAnonymous 215963.1445
5 recently rejected 211718.1931
6 Proof of Alpha 211712.3699
7 Cornfield Chase 210201.3605
8 Yan Su 210016.3629
9 Warkop PuteraPrakoso 201094.4945
10 iMak AI Lab 196481.1593
11 DKYnumber1 196323.2757
12 Quantumonster 193044.3011
13 S2WISH 189138.3673
14 Neko's Life 188642.9808
15 Autonomous Alpha 185936.6020
16 Lumia 185889.4276
17 Paragon 184046.5937
18 Made in Heaven 181871.6467
19 Money Miner 179846.1354
20 Verifiable Capital 178103.6278
21 PA-Agent 177343.4124
22 Jin & Pei 176009.6963
23 PumpkinChicken 164722.2406
24 Yongchang 161575.8191
25 Probably Right 157361.2474

T3 Simulation: teams 26–48 of 48

Rank Team Score
26 HackStreet Boys 156661.8792
27 fiftyfifty 152138.2119
28 lingsio 149569.8194
29 huhudawang 139213.3612
30 1991 130999.3163
31 chilli 121630.3250
32 Saifuddin 118112.9705
33 LastDigitsOfPi 97519.7633
34 Quant Science Flow 34835.2849
35 iluss 19429.7240
36 Apex 16467.6082
37 CMCCGDYDDICT 15559.8991
38 Banana Kingdom 14180.6623
39 Just-in-Time 13813.8474
40 Auror 13708.7408
41 Orrery 12998.5100
42 Outsiders 12116.2977
— TEst — not scored
— thalachira — not scored
— Hertz — not scored
— RRRRRR — not scored
— Artifical General Invesment — not scored
— Sequoia — not scored

No team matches that search.

T4 Explainability

T4 Explainability: teams 1–25 of 37

Rank Team Score
1 Genshin, launch! 0.8401
2 huhudawang 0.6237
3 Yan Su 0.5857
4 Just-in-Time 0.5787
5 MALIU 0.5592
6 Proof of Alpha 0.5381
7 Saifuddin 0.5234
8 abc123 0.5231
9 Paragon 0.4920
10 stone stone 0.4902
11 lingsio 0.4893
12 TradeFlare 0.4807
13 Sitadel Insecurities 0.4797
14 Northline 0.4797
15 yohoho 0.4696
16 Quiet Signal 0.4683
17 Winner 0.4520
18 WEI-SPBU 0.4371
19 OmniSync 0.4365
20 iluss 0.4314
21 VerityLovityFalsityCruelty 0.4278
22 Flowfront 0.4094
23 Pluto 0.4091
24 FAWDA 0.4074
25 ccczy 0.4029

T4 Explainability: teams 26–37 of 37

Rank Team Score
26 SignalCraft 0.3993
27 Chaewon-Research 0.3971
28 Apex 0.3954
29 1991 0.3739
30 DeltaNLP 0.3361
31 Quant Science Flow 0.3339
32 Richard Zhu 0.3319
33 scu_statistics 0.2413
34 LOOOONG 0.1475
35 Orrery -0.1788
— RRRRRR — not scored
— AAAA — not scored

No team matches that search.

2026 phases

Development, then Final + Verification.

Development runs through 12 October 2026. The joint Final + Verification phase runs from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE.

Upcoming

Development

Build and test with the public starter packages. Check the logged-in website for submission opening status.

0% complete
  1. Upcoming Aug 28 - Oct 12

    Development

    Teams practice and iterate with the public starter packages. Submission opening is announced separately.

  2. Upcoming Oct 13 - Oct 25

    Final + Verification

    One submission per entered track is evaluated on sealed private-test units. Organizers rerun leading submissions and review reproducibility in the same phase. There is no separate verification submission. Run outputs, logs, and score files stay hidden.

Awards

Cash awards in all four tracks.

The top three teams in each track win a cash award, decided by the final standings after verification. That is $3,000 per track and $12,000 in total.

  1. 1st place

    $1,500

    In each of the four tracks.

  2. 2nd place

    $1,000

    In each of the four tracks.

  3. 3rd place

    $500

    In each of the four tracks.

Amounts are in US dollars. A team's award is split equally among its registered members. Eligibility, verification, tax and payment requirements are set by the official rules, and a confirmed winner who accepts an award publishes the code needed to reproduce the method.

Scientific outputs

Agenthon is designed to explain failures, not just rank winners.

Cross-track failure map

Gate failures are labeled and aggregated to show where finance agents break down.

Information uplift

T2 investigates whether text and reasoning improve forecasts; this is not a separate scored component.

Speed-realism frontier

T3 measures throughput only after semantic fidelity and stylized facts survive checks.

Faithfulness under embargo

T4 requires evidence-backed predictions that do not cite future or unsupported facts.

Call for papers

The call for papers is closed.

Accepted papers will be presented as posters at the Agenthon workshop at NeurIPS in Atlanta on Saturday 12 December. The program will follow.

Our Supporters

Sponsor Agenthon 2026.

A number of sponsorship tiers and opportunities are available, including monetary and awards sponsorship, event space, data, infrastructure, and compute.

Agenthon Questions and Tasks are of real relevance to hedge funds and asset managers, spanning alpha forecasting, portfolio optimization, computational statistics, machine learning, and AI. Questions are provided jointly by the SQA, CEWIT and our Question Partners.

See Questions 2025 for last year's questions and a feel for Agenthon priorities.

Organizing Committee

Thank you from the organizers.

Lead Organizers

Pawel Polak

Assistant Professor, Department of Applied Mathematics and Statistics, Stony Brook University
Vice President, Society of Quantitative Analysts

Website · LinkedIn

Christos Koutsoyannis

Chief Investment Officer, Atlas Ridge Capital
Adjunct Professor, NYU Courant
Executive Advisory Board, Columbia Business School, Program for Financial Studies

Website · LinkedIn

Industry Co-Organizers

David Rosenberg

Head of Machine Learning Strategy, CTO Office
Bloomberg, Toronto, Canada

LinkedIn

Gary Kazantsev

Head of Quant Technology Strategy, Office of the CTO
Bloomberg, New York, USA

LinkedIn

Ioana Boier

Global Head of Capital Markets Strategy
NVIDIA Corporation, USA

LinkedIn

Track Leads

T1 · Coding

Quant-finance coding agents

Zhikang Dong
Track Lead T1
Independent Researcher

LinkedIn · GitHub

T2 · Forecasting

Reasoning-augmented time series

Ruolan Sun
Track Lead T2
Ph.D. Student, Stony Brook University

LinkedIn · GitHub

T3 · Simulation

Accelerated market simulation

Haohan Xu
Track Lead T3
Ph.D. Student, Stony Brook University

LinkedIn · GitHub

T4 · Explainability

Evidence-grounded prediction

Mathew Thiel
Track Lead T4
Quant Research Analyst, validityBase

LinkedIn · GitHub