AIDA.
Back to AIDAPrivate beta
Benchmarks

Real questions, real model calls, scored against independent SQL.

Every number on this page is read from recorded benchmark runs. Only cost is derived: it uses Groq list prices and the token counts each run reported.

Generated 2026-09-13 17:58 UTC from artifacts/benchmark/*.json via scripts/report_benchmark.py

48.1%70.4%correct on the 27 questions every compared run answeredgpt-oss-120b · Previous pipelineQwen3.8 27B · AIDA 4 · prompt 1 · repair 0
30wrong or unsafe answers on those questionsSame runs as above
112/112answers with a correct plan that also returned the correct rowsAcross all 5 runs
10/10requests correctly refused or clarifiedQwen3.8 27B · AIDA 4 · prompt 1 · repair 0
Before and after

The same 27 questions across 3 runs.

Only questions scored by every run below are counted, so each bar covers identical questions. logistics:BL42 was a required refusal for the previous pipeline, which had no calculations, and is scored as an answer for AIDA 4.

gpt-oss-120b · Previous pipeline13/27 correct (48.1%) · 3 wrong or unsafe
gpt-oss-20b · AIDA 4 · prompt 1 · repair 014/27 correct (51.8%) · 1 wrong or unsafe
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 019/27 correct (70.4%) · 0 wrong or unsafe
Correct answersCorrect refusalsFalse refusalsWrong or unsafeOther errors
Model comparison

AIDA 4 models on the same 42-question selection.

Complete runs only. Choose a measure.

Supported questions answered correctlyHigher is better
gpt-oss-20b · AIDA 4 · prompt 1 · repair 056.3%
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 068.8%
Engineering

Latency, cost and the code around the model.

Accuracy is half the story. These measurements show where each answer's time and money go, how the deterministic engine performs with no model at all, and the gates every change passes.

Where the time goes in one answerAverage per question, with provider rate-limit waits removed
gpt-oss-20b · AIDA 4 · prompt 1 · repair 0median answer 1.54 s · p95 2.00 s · AIDA code and SQL 4.0 ms (0.3%) · 42 questions
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 0median answer 1.75 s · p95 2.38 s · AIDA code and SQL 4.4 ms (0.3%) · 42 questions
gpt-oss-20b · AIDA 4 · prompt 3 · repair 1median answer 2.64 s · p95 21.96 s · AIDA code and SQL 7.7 ms (0.4%) · 9 questions
Qwen3.8 27B · AIDA 4 · prompt 3 · repair 1median answer 2.13 s · p95 2.40 s · AIDA code and SQL 8.7 ms (0.4%) · 6 questions
Prompt Guard screenModel calls (resolve, plan, repair)AIDA code and SQL
Cost per 1,000 questionsReported tokens at list price
gpt-oss-120b · Previous pipeline0.95 model calls per question$0.48
gpt-oss-20b · AIDA 4 · prompt 1 · repair 01.6 model calls per question · repair on 0% of questions$0.33
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 01.6 model calls per question · repair on 0% of questions$3.81
gpt-oss-20b · AIDA 4 · prompt 3 · repair 12.44 model calls per question · repair on 44% of questions$0.60
Qwen3.8 27B · AIDA 4 · prompt 3 · repair 12 model calls per question · repair on 0% of questions$5.44
Tokens per questionInput: instructions and approved catalog · Output: the plan
gpt-oss-120b · Previous pipeline2,517 tokens
gpt-oss-20b · AIDA 4 · prompt 1 · repair 03,469 tokens
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 03,420 tokens
gpt-oss-20b · AIDA 4 · prompt 3 · repair 16,502 tokens
Qwen3.8 27B · AIDA 4 · prompt 3 · repair 15,192 tokens
Input tokensOutput tokens
Time lost to provider rate limitsFree-tier waits, excluded from the answer times above
13 min39 of 42 questions waited · gpt-oss-20b · AIDA 4 · prompt 1 · repair 0
16 min41 of 42 questions waited · Qwen3.8 27B · AIDA 4 · prompt 1 · repair 0
46 min8 of 9 questions waited · gpt-oss-20b · AIDA 4 · prompt 3 · repair 1
3 min5 of 6 questions waited · Qwen3.8 27B · AIDA 4 · prompt 3 · repair 1
Deterministic engine, no modelEvery benchmark question with an expected plan, executed 5 times with result caches cleared
129/129explicit plans returned the independent oracle rows
2.3 msmedian uncached plan, request to result (p95 8.6 ms)
2.2 msmedian time inside the database
0.1 msmedian repeat served from the result cache (129 of 129)
SuitePlansRows match oracleUncached p50Uncached p95Cached p50
Logistics regression4040/405.5 ms13 ms0.2 ms
Relational demos3131/311.3 ms2.5 ms0.1 ms
Name resolution33/30.5 ms1.0 ms0.1 ms
Semantic regression5555/551.4 ms4.7 ms0.1 ms
Recorded model outputs replayed through today's codeNo model calls: validation, SQL and calculations only
RunAnswers replayedIdentical rowsIdentical planCode and SQL p50Code and SQL p95Repeat question p50
gpt-oss-20b · AIDA 4 · prompt 1 · repair 02020/2020/204.3 ms13 ms0.4 ms20/20 with no model call
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 02222/2222/224.4 ms13 ms0.6 ms22/22 with no model call
gpt-oss-20b · AIDA 4 · prompt 3 · repair 188/88/86.3 ms13 ms0.6 ms8/8 with no model call
Qwen3.8 27B · AIDA 4 · prompt 3 · repair 166/66/67.7 ms13 ms0.6 ms6/6 with no model call
Quality gatesRecorded 2026-09-13
323backend tests passed
32attack and misuse tests
34interpreter contract tests, including malformed model output
9/9browser journey steps passed (2026-09-13)
Configured protectionsRead from the running code
Sign-in attempts per client20 per minute
Sign-in attempts per email5 per 15 min
Sign-ups per client5 per hour
Questions per user20 per minute
Explicit plans per user120 per minute
Model-backed questions, all users60 per minute
Uploads per user5 per 10 min
Catalog changes per user20 per 10 min
Onboarding saves per user20 per 10 min
Misuse pause5 restricted refusals in 10 min pause questions for 15 min

Engine and replay figures come from scripts/benchmark_engine.py on Windows, Python 3.13.15, 16 logical CPUs. Model latency and cost come from the recorded runs on this page.

Where answers were lost

What stopped each run on the questions it missed.

A false refusal is recorded with the reason the model gave, or as rejected by code checks when the model's reply broke AIDA's grounding or format rules.

gpt-oss-20b · AIDA 4 · prompt 1 · repair 015 of 42 questions missed
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 010 of 42 questions missed
Model asked to disambiguateReply rejected by code checksModel did not recognise a nameModel judged it unsupportedModel called it too vagueWrong or unsafe answer
When the plan was right, were the rows right?Answers whose plan matched the expected plan, and how many of those returned the expected rows
58/58gpt-oss-120b · Previous pipeline
18/18gpt-oss-20b · AIDA 4 · prompt 1 · repair 0
22/22Qwen3.8 27B · AIDA 4 · prompt 1 · repair 0
8/8gpt-oss-20b · AIDA 4 · prompt 3 · repair 1
6/6Qwen3.8 27B · AIDA 4 · prompt 3 · repair 1
By suite

Qwen3.8 27B · AIDA 4 · prompt 1 · repair 0 by kind of question.

Calculations
Answers 3/3
Logistics regression
Answers 12/18
Refusals 4/4
Relational demos
Answers 3/5
Refusals 1/1
Name resolution
Answers 1/1
Refusals 2/2
Semantic regression
Answers 3/5
Refusals 3/3
Prompt and repair changes

Same model, updated prompt, on the questions both runs scored.

These runs are partial: the provider's daily token quota stopped them. Small question counts are shown as they are.

gpt-oss-20b9 shared questions
AIDA 4 · prompt 1 · repair 0Complete6/9
AIDA 4 · prompt 3 · repair 1Partial8/9
Qwen3.8 27B6 shared questions
AIDA 4 · prompt 1 · repair 0Complete5/6
AIDA 4 · prompt 3 · repair 1Partial · stopped by provider quota6/6
All runs

Every recorded run.

RunStatusQuestions scoredSupported correctRefusals correctWrong or unsafeFalse refusalsTokens / questionMedian answerCost / 100 questions
gpt-oss-120b · Previous pipelinebaseline-current-gpt-oss-120bPartial · stopped by provider quota8458/71 81.7%11/13 84.6%3122,517Not recorded$0.048
gpt-oss-20b · AIDA 4 · prompt 1 · repair 0aida4-two_stage-gpt-oss-20b-selectionComplete4218/32 56.3%9/10 90.0%2133,4691.54 s$0.033
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 0aida4-two_stage-qwen3.8-27b-selectionComplete4222/32 68.8%10/10 100.0%0103,4201.75 s$0.381
gpt-oss-20b · AIDA 4 · prompt 3 · repair 1aida4v3-two_stage-gpt-oss-20b-selectionPartial9 (+1 not run)8/9 88.9%0/0 016,5022.64 s$0.060
Qwen3.8 27B · AIDA 4 · prompt 3 · repair 1aida4v3-two_stage-qwen3.8-27b-selectionPartial · stopped by provider quota66/6 100.0%0/0 005,1922.13 s$0.544

Runs with different question sets are not directly comparable; use the before-and-after and model comparison sections for like-for-like figures. The previous-pipeline runner did not record time spent waiting on provider rate limits, so its answer time is not shown.

Every question

Search, filter and sort the 99 questions and their outcomes.

99 rows
logistics:BL01Logistics regression (joins, subqueries, archives, refusals)Logistics sampleOur network planning sheet needs a current leg total. How many shipment legs are recorded?AnswerCorrectCorrectCorrectCorrectCorrect
logistics:BL02Logistics regression (joins, subqueries, archives, refusals)Logistics sampleSome consignments travel on several legs. How many different consignments are represented in the current leg records?AnswerCorrect
logistics:BL03Logistics regression (joins, subqueries, archives, refusals)Logistics sampleFor the current network activity, what is the total transported weight in kilograms?AnswerCorrect
logistics:BL04Logistics regression (joins, subqueries, archives, refusals)Logistics sampleHow much has the current leg activity incurred in handling charges, expressed in US dollars?AnswerCorrect
logistics:BL05Logistics regression (joins, subqueries, archives, refusals)Logistics sampleWhat is the average recorded transit time, in minutes per leg, across the current records?AnswerCorrect
logistics:BL06Logistics regression (joins, subqueries, archives, refusals)Logistics samplePrepare a customer-market breakdown of current handling charges and transported kilograms, with the largest charge total first.AnswerCorrectCorrectCorrectCorrectCorrect
logistics:BL07Logistics regression (joins, subqueries, archives, refusals)Logistics sampleFor each customer market and destination zone together, report current handling charges and the number of distinct consignments; put the largest charges first.AnswerFalse refusalCorrectCorrectCorrectCorrect
logistics:BL08Logistics regression (joins, subqueries, archives, refusals)Logistics sampleCompare current service-class and carrier combinations using average transit minutes and leg count. Start with the shortest average, breaking ties alphabetically by service class and then carrier.AnswerFalse refusalFalse refusalFalse refusalCorrectCorrect
logistics:BL09Logistics regression (joins, subqueries, archives, refusals)Logistics sampleHow are current legs distributed across shipment states? Keep unassigned states as a group and list the state labels alphabetically.AnswerCorrect
logistics:BL10Logistics regression (joins, subqueries, archives, refusals)Logistics sampleGive me the current transported kilograms month by month, starting with the earliest departure month.AnswerCorrect
logistics:BL11Logistics regression (joins, subqueries, archives, refusals)Logistics sampleFor current legs whose service tariff is more than 1,800 cents, what are the average transit minutes and total handling charges?AnswerCorrect
logistics:BL12Logistics regression (joins, subqueries, archives, refusals)Logistics sampleCount the current legs whose whole-consignment charge is at least 20,000 cents but whose own handling charge is below 5,000 cents, and give their total handling charges in dollars.AnswerCorrectCorrectCorrectFalse refusalCorrect
logistics:BL13Logistics regression (joins, subqueries, archives, refusals)Logistics sampleIn the Retail customer market, how many current legs carry exactly four packages?AnswerFalse refusalFalse refusalCorrectCorrectCorrect
logistics:BL14Logistics regression (joins, subqueries, archives, refusals)Logistics sampleWhat are the current transported weight and handling charges for Express or Overnight legs going to North or West destination zones, when route distance is at least 1,000 kilometers?AnswerCorrect
logistics:BL15Logistics regression (joins, subqueries, archives, refusals)Logistics sampleFor current legs marked Delayed whose shipment state is Delivered, show the leg count and average transit minutes.AnswerCorrect
logistics:BL16Logistics regression (joins, subqueries, archives, refusals)Logistics sampleFor active carriers, count current legs in every service class except Freight, split by carrier and listed alphabetically.AnswerFalse refusalCorrectFalse refusalCorrect
logistics:BL17Logistics regression (joins, subqueries, archives, refusals)Logistics sampleWhich destination zones have more than 850,000 kilograms of current transported weight? Show each qualifying total, largest first.AnswerCorrectCorrectCorrectCorrect
logistics:BL18Logistics regression (joins, subqueries, archives, refusals)Logistics sampleOnly consider current legs with a service tariff above 1,000 cents. Among their carriers, show those whose total handling charges exceed $73,000, highest total first.AnswerCorrect
logistics:BL19Logistics regression (joins, subqueries, archives, refusals)Logistics sampleList service classes represented by at least 2,000 distinct consignments in current legs, showing those distinct counts from highest to lowest.AnswerCorrect
logistics:BL20Logistics regression (joins, subqueries, archives, refusals)Logistics sampleHow many current legs have at least one shipment-leg exception recorded?AnswerFalse refusalFalse refusalCorrectCorrect
logistics:BL21Logistics regression (joins, subqueries, archives, refusals)Logistics sampleFor current legs with no recorded exceptions at all, give the transported weight and handling-charge totals.AnswerCorrect
logistics:BL22Logistics regression (joins, subqueries, archives, refusals)Logistics sampleHow much handling charge did current legs incur when they had a Damage exception of severity at least 3? Also count those legs once each.AnswerCorrect
logistics:BL23Logistics regression (joins, subqueries, archives, refusals)Logistics sampleCount current legs for which no exception has been marked Resolved; legs without any exception should count too.AnswerCorrect
logistics:BL24Logistics regression (joins, subqueries, archives, refusals)Logistics sampleGive the current leg count and transported kilograms after excluding any leg with an Open Weather exception. Other exceptions may remain.AnswerFalse refusalFalse refusalCorrectNot run (quota)
logistics:BL25Logistics regression (joins, subqueries, archives, refusals)Logistics sampleFor active carriers, compare average transit minutes and leg counts among current legs with a Weather or Customs exception of severity at least 4. Put the longest average first.AnswerFalse refusalFalse refusalFalse refusal
Page 1 of 4
Method and limits

How these numbers were produced.

  • Each question runs through the same code path as the product, with real model inference, Prompt Guard screening, validation, SQL compilation and calculations.
  • Expected rows come from independently written SQL or independently calculated values. Expected plans and rows are never sent to the model.
  • A supported question counts as correct only when both the plan and the rows match. A refusal case counts as correct only when AIDA does not answer.
  • These questions are a regression set: prompts were improved after inspecting failures, so these are not blind results on unseen data.
  • Answer times exclude time spent waiting on provider rate limits. Costs are estimates from reported token usage and Groq list prices.
  • Runs marked partial were stopped by the provider's daily token quota and cover fewer questions.
  • No AIDA 4 results are recorded yet for gpt-oss-120b.

See the answers for yourself.

Every answer in AIDA shows its plan, SQL and lineage.