#### Machine learning

### [Planning is unsolved](/content/blog/planning-is-unsolved/index.html)
May 4, 2026  
Models are getting better at long-horizon tasks. So why don't they help much when you're planning clinical development or thinking through a launch?

### [Against RL: The Case for System 2 Learning](/content/blog/system-2-learning/index.html)
Jan 30, 2025  
Reinforcement learning may boost LLMs today, but it cannot deliver safe, long-term intelligence. We argue for System 2 learning instead.

### [Trust at scale: Auto-evaluation for high-stakes LLM accuracy](/content/blog/auto-evaluation/index.html)
Jul 23, 2024  
Elicit develops LLM-based auto-evals to balance scale, trust, and flexibility, ensuring reliable scientific reasoning at superhuman speed.

### [Factored Verification: Detecting and Reducing Hallucinations in Frontier Models Using AI Supervision](/content/blog/factored-verification-detecting-and-reducing-hallucinations-in-frontier-models-using-ai-supervision/index.html)
Oct 20, 2023  
We evaluate an automated approach for catching hallucinations in paper abstracts, aiming for consistently more trustworthy results.
