OlympicRAG Benchmark
TigerGraph Agentic GraphRAG Hackathon · Round 1 benchmark

OlympicRAG Benchmark

Scorecard

100 public questions, each answered by all three pipelines on the same TigerGraph graph and the same LLM. Accuracy is judged against the ground truth; citation recall measures how many gold documents the answer cites.

Grounding check

Every answer is checked in code against the evidence it cites. It counts as supported only if all of it comes from one cited event or source. An answer fails if it combines facts from different events, contains a name or number no cited evidence states, or names a non-gold medallist as the winner. This needs no ground truth, so it also covers the hidden set.

By question type

Same questions, split by what they ask for. The left chart shows how often each pipeline is right; the right shows what it spends per question.

Accuracy (% correct)

Tokens per question

When does the agent pay off?

For each question type, the agent is compared with the cheaper of the two single-pass pipelines that scores best. A gain of at least 10 points marks the type as one that needs an agent.

Harder multi-step questions (our own set)

36 questions built from the same corpus that need more than one step: two editions back (the 1994 Winter Games break the four-year pattern), venue + date then the previous edition, an athlete's gold count, comparing two events, the next Games, a silver medallist, and five events whose articles have no infobox, so the graph has nothing and the agent must switch to reading the article. Every gold answer is computed by code from the parsed infoboxes or read from the article text with the quote recorded; none is written by an LLM.

Agent paths on the harder set

Questions outside the graph: films (our own set)

The corpus also holds 546 film articles that are not in the graph. These 12 questions describe a film without naming it (year, director, a star), so the graph has nothing and title matching cannot find it: the article has to be found by meaning with similarity search, then read. Gold answers come from the film infobox, taken by code.

Agent paths on these questions

Reasoning over time: conflicting and superseded facts

Some facts in the corpus have more than one version. A medallist "originally won" and was later disqualified, or the infobox and the article text give different competitor counts. Ingestion stores each version as a Claim vertex in TigerGraph, linked to its event and to the claim it supersedes (code only, no LLM). The fact_history agent reads them, and the harness writes a note that names both versions. 18 questions test this: the original winner, the current one, the year of the decision, a medal left vacant, and five count conflicts where the official record must be used and the disagreement reported.

How the agent investigates

Which specialised agents the orchestrator called, how often they returned evidence, the routes it took, and why it stopped.

Specialised agents invoked

Most common investigation paths

Question explorer

Every public question with all three answers. Select a row to open the agent's full trace: each decision, the tool it called, time and tokens per step, and what came back.

Hidden set (50 questions)

Answers, token counts and full agent traces for the 50 held-out questions are submitted as raw outputs in results/hidden_*.jsonl; the organisers score them.