Projects
LLM reliability

TrueAI

Detecting LLM hallucinations, plus 150+ customer discovery interviews.

Hallucination detectionEvaluationNSF I-Corps

What it is

A project on detecting when large language models make things up (February to May 2025). I built an evaluation harness that scores whether an LLM’s answer is actually grounded in its source.

0.85AUC for our approach
0.79AUC for SAPLMA
0.62AUC for SelfCheckGPT
<100 mstarget latency for real-time scoring

Measured on annotated benchmarks.

Customer discovery

I took part in 150+ NSF I-Corps interviews with AI leaders, engineers, researchers and product managers. A key finding: enterprises need fast, low-latency guardrails, not slow "LLM-as-a-judge" re-checks, which redirected the design from LLM-as-a-judge re-checks to sub-100 ms scoring based on the model’s own logits and entropy.

All projects Ask Gizmo