Frontier models can recover up to 65% of facts they can't directly recall — just by thinking longer

EDITOR BRIEF
Researchers at Google Research and Technion found that frontier models such as GPT-5 and Gemini-3 may encode 95% to 98% of tested facts, even when they fail to recall them in normal prompting. The study argues many hallucinations stem from recall failures, not missing knowledge, and proposes fact-level profiling to distinguish stored facts from reliably usable knowledge.
INSIGHTS
The findings suggest AI reliability may improve through better inference-time techniques rather than only bigger models, more training data, or retrieval systems. If validated broadly, test-time compute could become a key lever for factual accuracy, reshaping how teams diagnose and mitigate hallucinations.
COMMENTS
Discussion
> geekhaus:~$ next read?
Next read recommendations

VentureBeat
Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities

VentureBeat
Meta prices Muse Voice Transcribe at $0.18 an hour, with real-time diarization for 20+ speakers: a steal for enterprises?
TechCrunch