← LinkedIn posts

3× the accuracy on business questions, from the same model

Researchers put 43 business questions to GPT-4 over an industry-standard insurance data model. What moved accuracy from 16.7% to 54.2% was not the model. It was what the model was given to read.

GPT-4 answering 43 business questions over an insurance data model: 16.7% correct from the raw database, and none correct when the answer spanned more than four tables, against 54.2% with a knowledge graph and over a third correct on those same hard questions. Beside it, the follow-up studies that isolate business context from model choice, and an ontology check scoring 72% correct, 8% "I don't know" and 20% wrong.
The whole post in one graphic. Open the graphic full size ↗

3× the accuracy on business questions, from the same AI model. What changed was what it was given to read.

Why it matters

Researchers asked GPT-4 43 business questions over an industry-standard insurance data model. Working from the raw database, it got 1 in 6 right on average.

Then they gave it a knowledge graph: a map of the business itself. Claims, policies, premiums, agents, and how each one connects.

Accuracy rose from 16.7% to 54.2%.

The gap was widest on hard questions. When an answer spanned more than 4 tables, the raw-database setup got none right. With the map, it got over a third.

Before asking how smart the AI is, ask what it knows about the business.

How it works

Sequeda, Allemang, and Jacob, who built the benchmark, raise the fair objection themselves: GPT-4 also switched from SQL to SPARQL, so the map was not the only thing that changed.

A 2026 paired study by Cube, a semantic-layer vendor, isolates the variable. SQL in both conditions, with only a 4 KB business-context document varying. Claude Opus 4.7, Claude Sonnet 4.6, and GPT-5.4 each gained 17 to 23 points and, with the document in hand, landed within one point of each other. In that study, context mattered far more than model choice. The vendor sells semantic layers, so the framing deserves the usual discount; the design is what makes it worth reading, because only one thing moved.

A follow-up by Allemang and Sequeda added query checks against the ontology: 72% correct, 8% “I don’t know”, 20% wrong. An honest “I don’t know” beats a confident wrong number, because the confident wrong number is the one that reaches a slide.

Two limits are worth stating plainly. This graph used RDF and SPARQL, so the result says nothing about any specific graph database. And someone has to build the map and keep it current: the accuracy arrives with a maintenance bill attached, and that bill is the part most pilots forget to price.