LinkedIn posts
The systems I build, and the research behind them
Each post takes one decision out of a system I built or a proof I worked through: what broke, what I changed, and what the number was afterward.
-
47 to 38: how far a graph-backed AI fell when facts went missing
A peer-reviewed study removed facts from a knowledge graph on purpose. With 4 in 10 of the needed facts gone, the graph-backed system scored what the model scored with no graph at all.
-
The costlier AI search setup lost on simple fact questions
A peer-reviewed KDD 2026 study ran plain retrieval and graph retrieval through the same tests. Which one won depended on the question, and the AI judge scoring the contest sometimes reversed its verdict when the answers swapped places.
-
3× the accuracy on business questions, from the same model
Researchers put 43 business questions to GPT-4 over an industry-standard insurance data model. What moved accuracy from 16.7% to 54.2% was not the model. It was what the model was given to read.
-
The safety net answered 32 of 150 questions from the wrong document
If the search found nothing in the right document, my system searched everything, so a user always got an answer. That fallback turned a visible failure into a confident wrong answer.
-
The agent ran the checks, and wrote the report saying they passed
Gatekeepers answer whether an action was allowed. They do not answer whether the agent did what it says it did. A session-end hook re-runs seven of the eight gates, so the agent's summary stops being the record.
-
I gave three coding agents less capability on purpose
Claude Code ships a general-purpose agent that can do almost anything in a codebase. I run three specialists beside it, with 46 explicit prohibitions between them, because when mistakes cost money capability stops being the constraint.
-
Five quality gates passed, graded by the agent that wrote the code
The agent wrote the code, wrote its tests, ran the checks, and reported five green. Nothing in the pipeline was built to ask who checked the checker.
-
Do I have a guarantee, or am I guessing?
Convergence analysis taught me to ask one question of any algorithm. That question shaped an agentic supply-chain system built to decide when not to trust its own answers.
-
A hand-tuned prompt scored 91%. A compiled one scored 99%.
Route a supply-chain question wrong and an inventory query comes back with a vehicle-routing answer. I replaced the hand-written classifier prompt with a DSPy-compiled one, then spent longer on the measurement than on the compile.
-
Same question, different numbers, wrong answer
"Allocate 400 units" and "allocate 1,000 units" embed almost identically. A cache that matches on similarity alone will serve one manager another manager's numbers.
-
Why we proved convergence on the simplest network there is
The linear layer is the one component every neural network shares. Its loss is still nonconvex, and training it is still NP-complete. That is why it is the right place to look for a guarantee.
-
I defended my PhD: convergence analysis of training deep linear neural networks
The learning rate depends on constants existing theory can't compute before training, so they get chosen by trial and error. I closed part of that gap, and I'm now applying the same rigor to production AI.
-
The copilot keeps answering when the language model goes dark
A free tier capped at 50 API calls a day forced every design decision in this system. The constraint is gone now. The architecture it produced stayed, because it earned its place.
-
The most reliable part of the copilot is the part the LLM never touches
Ask a language model to plan a logistics decision and you get the most statistically likely answer, not the cheapest one. So I kept the model out of the math and spent the effort on the boundary between them.
-
$10,000 is where the agent stops and a human decides
Every approval workflow needs a number. Below it the agent executes, above it a person signs off. The bug that taught me where to put that check only surfaced in integration testing.