Real questions, real model calls, scored against independent SQL.
Every number on this page is read from recorded benchmark runs. Only cost is derived: it uses Groq list prices and the token counts each run reported.
The same 27 questions across 3 runs.
Only questions scored by every run below are counted, so each bar covers identical questions. logistics:BL42 was a required refusal for the previous pipeline, which had no calculations, and is scored as an answer for AIDA 4.
AIDA 4 models on the same 42-question selection.
Complete runs only. Choose a measure.
Latency, cost and the code around the model.
Accuracy is half the story. These measurements show where each answer's time and money go, how the deterministic engine performs with no model at all, and the gates every change passes.
| Suite | Plans | Rows match oracle | Uncached p50 | Uncached p95 | Cached p50 |
|---|---|---|---|---|---|
| Logistics regression | 40 | 40/40 | 5.5 ms | 13 ms | 0.2 ms |
| Relational demos | 31 | 31/31 | 1.3 ms | 2.5 ms | 0.1 ms |
| Name resolution | 3 | 3/3 | 0.5 ms | 1.0 ms | 0.1 ms |
| Semantic regression | 55 | 55/55 | 1.4 ms | 4.7 ms | 0.1 ms |
| Run | Answers replayed | Identical rows | Identical plan | Code and SQL p50 | Code and SQL p95 | Repeat question p50 |
|---|---|---|---|---|---|---|
| gpt-oss-20b · AIDA 4 · prompt 1 · repair 0 | 20 | 20/20 | 20/20 | 4.3 ms | 13 ms | 0.4 ms20/20 with no model call |
| Qwen3.8 27B · AIDA 4 · prompt 1 · repair 0 | 22 | 22/22 | 22/22 | 4.4 ms | 13 ms | 0.6 ms22/22 with no model call |
| gpt-oss-20b · AIDA 4 · prompt 3 · repair 1 | 8 | 8/8 | 8/8 | 6.3 ms | 13 ms | 0.6 ms8/8 with no model call |
| Qwen3.8 27B · AIDA 4 · prompt 3 · repair 1 | 6 | 6/6 | 6/6 | 7.7 ms | 13 ms | 0.6 ms6/6 with no model call |
| Sign-in attempts per client | 20 per minute |
| Sign-in attempts per email | 5 per 15 min |
| Sign-ups per client | 5 per hour |
| Questions per user | 20 per minute |
| Explicit plans per user | 120 per minute |
| Model-backed questions, all users | 60 per minute |
| Uploads per user | 5 per 10 min |
| Catalog changes per user | 20 per 10 min |
| Onboarding saves per user | 20 per 10 min |
| Misuse pause | 5 restricted refusals in 10 min pause questions for 15 min |
Engine and replay figures come from scripts/benchmark_engine.py on Windows, Python 3.13.15, 16 logical CPUs. Model latency and cost come from the recorded runs on this page.
What stopped each run on the questions it missed.
A false refusal is recorded with the reason the model gave, or as rejected by code checks when the model's reply broke AIDA's grounding or format rules.
Qwen3.8 27B · AIDA 4 · prompt 1 · repair 0 by kind of question.
Same model, updated prompt, on the questions both runs scored.
These runs are partial: the provider's daily token quota stopped them. Small question counts are shown as they are.
Every recorded run.
| Run | Status | Questions scored | Supported correct | Refusals correct | Wrong or unsafe | False refusals | Tokens / question | Median answer | Cost / 100 questions |
|---|---|---|---|---|---|---|---|---|---|
| gpt-oss-120b · Previous pipelinebaseline-current-gpt-oss-120b | Partial · stopped by provider quota | 84 | 58/71 81.7% | 11/13 84.6% | 3 | 12 | 2,517 | Not recorded | $0.048 |
| gpt-oss-20b · AIDA 4 · prompt 1 · repair 0aida4-two_stage-gpt-oss-20b-selection | Complete | 42 | 18/32 56.3% | 9/10 90.0% | 2 | 13 | 3,469 | 1.54 s | $0.033 |
| Qwen3.8 27B · AIDA 4 · prompt 1 · repair 0aida4-two_stage-qwen3.8-27b-selection | Complete | 42 | 22/32 68.8% | 10/10 100.0% | 0 | 10 | 3,420 | 1.75 s | $0.381 |
| gpt-oss-20b · AIDA 4 · prompt 3 · repair 1aida4v3-two_stage-gpt-oss-20b-selection | Partial | 9 (+1 not run) | 8/9 88.9% | 0/0 — | 0 | 1 | 6,502 | 2.64 s | $0.060 |
| Qwen3.8 27B · AIDA 4 · prompt 3 · repair 1aida4v3-two_stage-qwen3.8-27b-selection | Partial · stopped by provider quota | 6 | 6/6 100.0% | 0/0 — | 0 | 0 | 5,192 | 2.13 s | $0.544 |
Runs with different question sets are not directly comparable; use the before-and-after and model comparison sections for like-for-like figures. The previous-pipeline runner did not record time spent waiting on provider rate limits, so its answer time is not shown.
Search, filter and sort the 99 questions and their outcomes.
| logistics:BL01 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Our network planning sheet needs a current leg total. How many shipment legs are recorded? | Answer | Correct | Correct | Correct | Correct | Correct |
| logistics:BL02 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Some consignments travel on several legs. How many different consignments are represented in the current leg records? | Answer | Correct | — | — | — | — |
| logistics:BL03 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | For the current network activity, what is the total transported weight in kilograms? | Answer | Correct | — | — | — | — |
| logistics:BL04 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | How much has the current leg activity incurred in handling charges, expressed in US dollars? | Answer | Correct | — | — | — | — |
| logistics:BL05 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | What is the average recorded transit time, in minutes per leg, across the current records? | Answer | Correct | — | — | — | — |
| logistics:BL06 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Prepare a customer-market breakdown of current handling charges and transported kilograms, with the largest charge total first. | Answer | Correct | Correct | Correct | Correct | Correct |
| logistics:BL07 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | For each customer market and destination zone together, report current handling charges and the number of distinct consignments; put the largest charges first. | Answer | False refusal | Correct | Correct | Correct | Correct |
| logistics:BL08 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Compare current service-class and carrier combinations using average transit minutes and leg count. Start with the shortest average, breaking ties alphabetically by service class and then carrier. | Answer | False refusal | False refusal | False refusal | Correct | Correct |
| logistics:BL09 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | How are current legs distributed across shipment states? Keep unassigned states as a group and list the state labels alphabetically. | Answer | Correct | — | — | — | — |
| logistics:BL10 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Give me the current transported kilograms month by month, starting with the earliest departure month. | Answer | Correct | — | — | — | — |
| logistics:BL11 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | For current legs whose service tariff is more than 1,800 cents, what are the average transit minutes and total handling charges? | Answer | Correct | — | — | — | — |
| logistics:BL12 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Count the current legs whose whole-consignment charge is at least 20,000 cents but whose own handling charge is below 5,000 cents, and give their total handling charges in dollars. | Answer | Correct | Correct | Correct | False refusal | Correct |
| logistics:BL13 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | In the Retail customer market, how many current legs carry exactly four packages? | Answer | False refusal | False refusal | Correct | Correct | Correct |
| logistics:BL14 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | What are the current transported weight and handling charges for Express or Overnight legs going to North or West destination zones, when route distance is at least 1,000 kilometers? | Answer | Correct | — | — | — | — |
| logistics:BL15 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | For current legs marked Delayed whose shipment state is Delivered, show the leg count and average transit minutes. | Answer | Correct | — | — | — | — |
| logistics:BL16 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | For active carriers, count current legs in every service class except Freight, split by carrier and listed alphabetically. | Answer | False refusal | Correct | False refusal | Correct | — |
| logistics:BL17 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Which destination zones have more than 850,000 kilograms of current transported weight? Show each qualifying total, largest first. | Answer | Correct | Correct | Correct | Correct | — |
| logistics:BL18 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Only consider current legs with a service tariff above 1,000 cents. Among their carriers, show those whose total handling charges exceed $73,000, highest total first. | Answer | Correct | — | — | — | — |
| logistics:BL19 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | List service classes represented by at least 2,000 distinct consignments in current legs, showing those distinct counts from highest to lowest. | Answer | Correct | — | — | — | — |
| logistics:BL20 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | How many current legs have at least one shipment-leg exception recorded? | Answer | False refusal | False refusal | Correct | Correct | — |
| logistics:BL21 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | For current legs with no recorded exceptions at all, give the transported weight and handling-charge totals. | Answer | Correct | — | — | — | — |
| logistics:BL22 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | How much handling charge did current legs incur when they had a Damage exception of severity at least 3? Also count those legs once each. | Answer | Correct | — | — | — | — |
| logistics:BL23 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Count current legs for which no exception has been marked Resolved; legs without any exception should count too. | Answer | Correct | — | — | — | — |
| logistics:BL24 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | Give the current leg count and transported kilograms after excluding any leg with an Open Weather exception. Other exceptions may remain. | Answer | False refusal | False refusal | Correct | Not run (quota) | — |
| logistics:BL25 | Logistics regression (joins, subqueries, archives, refusals) | Logistics sample | For active carriers, compare average transit minutes and leg counts among current legs with a Weather or Customs exception of severity at least 4. Put the longest average first. | Answer | False refusal | False refusal | False refusal | — | — |
How these numbers were produced.
- Each question runs through the same code path as the product, with real model inference, Prompt Guard screening, validation, SQL compilation and calculations.
- Expected rows come from independently written SQL or independently calculated values. Expected plans and rows are never sent to the model.
- A supported question counts as correct only when both the plan and the rows match. A refusal case counts as correct only when AIDA does not answer.
- These questions are a regression set: prompts were improved after inspecting failures, so these are not blind results on unseen data.
- Answer times exclude time spent waiting on provider rate limits. Costs are estimates from reported token usage and Groq list prices.
- Runs marked partial were stopped by the provider's daily token quota and cover fewer questions.
- No AIDA 4 results are recorded yet for gpt-oss-120b.
See the answers for yourself.
Every answer in AIDA shows its plan, SQL and lineage.