Frontier models can recover up to 65% of facts they can't directly recall — just by thinking longer

EDITOR BRIEF
Researchers at Google Research and Technion found that frontier models such as GPT-5 and Gemini-3 may encode 95% to 98% of tested facts, even when they fail to recall them in normal prompting. The study argues many hallucinations stem from recall failures, not missing knowledge, and proposes fact-level profiling to distinguish stored facts from reliably usable knowledge.
INSIGHTS
The findings suggest AI reliability may improve through better inference-time techniques rather than only bigger models, more training data, or retrieval systems. If validated broadly, test-time compute could become a key lever for factual accuracy, reshaping how teams diagnose and mitigate hallucinations.
COMMENTS
Discussion
> geekhaus:~$ next read?
Next read recommendations

VentureBeat
Closing an Azure OpenAI assistant's retrieval gap didn't take a new identity platform. It took one filter and a narrower assistant.

VentureBeat
Software engineers' new job isn't writing code — it's designing the boundaries AI agents can't break

VentureBeat