London AI lab Inherent says its Faraday agent outperforms Anthropic and OpenAI models in research replication
A modest model outperforms larger rivals
Inherent, a London‑based AI laboratory founded by former Google DeepMind researchers, has unveiled a new result in AI‑assisted science.
The startup claims its AI agent, named Faraday, reproduced the findings of published scientific papers more accurately than larger models from Anthropic and OpenAI.
The comparison used Anthropic’s Claude Opus 4.8 and OpenAI’s GPT‑5.5, both of which are considerably larger than Faraday’s underlying model.
Faraday runs on Qwen 3.6, a model with just 27 billion parameters, a fraction of the size of its rivals.
Parameter count serves as a proxy for model scale and training cost, making Faraday’s efficiency noteworthy.
Training for scientific taste
The task of independent paper replication is a standard exercise for human researchers, according to co‑founder and chief scientist Edward Hughes.
“Many PhD students actually start by doing this,” Hughes told TechCrunch.
Inherent framed the achievement not merely as a win over competitors but as evidence of a novel building approach.
“What was most interesting to us about this was not so much the result of beating those frontier agents — which of course we liked — but was actually the way we went about building this,” Hughes said.
The startup’s higher bar required Faraday to show “research taste,” an ability to judge which experiments are worthwhile and how to design them.
Teaching such intangible judgment relies on reinforcement learning, a method that rewards desirable outcomes instead of prescribing explicit rules.
Inherent prefers reward‑based training over direct instruction on the scientific method, hoping the approach will generalize across disciplines.
“We’re always guided by that north star of building an AI scientist agent and imbuing our agents with taste,” Hughes added.
Company philosophy and future direction
Rather than creating its own code‑generation tool, Faraday leverages OpenAI’s GPT‑5.5 Codex for programming tasks.
This mirrors how human scientists typically use existing software instead of building every component from scratch.
The company also aims to avoid agents that simply echo user expectations, seeking instead a collaborative partner.
Hughes described his ideal teammate as one who says, “I got curious about this, and I went off and I did these experiments. What do you think of these results?”
Inherent’s team of roughly a dozen engineers works together in person at a King’s Cross office.
The location reflects London’s emergence as a global AI hub, a transformation partly driven by DeepMind’s presence.
“We believe that London is the place to be,” Hughes remarked.
Hughes also voiced support for ending the UK practice of “garden leave,” which restricts departing employees from joining rivals for several months.
He noted that the restriction is uncommon in the United States, where many AI researchers operate.
Inherent raised a $50 million seed round just weeks before emerging from stealth mode.
The funding provides resources to expand Faraday’s capabilities beyond replication toward original scientific discovery.
By demonstrating that a 27‑billion‑parameter model can surpass larger systems on a rigorous benchmark, Inherent challenges the assumption that scale alone drives performance.
Investors may see the result as a signal that efficient, purpose‑built agents can compete with frontier‑scale models.
The achievement also offers a proof‑of‑concept for reinforcement‑learning‑based training of AI scientists.
If Faraday can later generate novel hypotheses, the impact on research productivity could be significant.
For now, the replication test serves as a concrete metric that validates Inherent’s training philosophy.
The startup’s next public milestones are expected to focus on expanding Faraday’s domain expertise and publishing peer‑reviewed results.
Why This Matters: Inherent’s Faraday shows that a modest‑size model can outpace larger rivals at scientific replication, suggesting a new efficiency path for AI‑driven research.
This digest was compiled from:
Share this digest
People Also Ask
- NITDA Announces Initiative to Standardize Innovation Hubs Nationwide
NITDA introduced a guide to standardize innovation hubs, aiming to spread tech support beyond Lagos and Abuja across Nigeria.
- Over a million users have pressed LinkedIn’s “AI slop” button
LinkedIn introduces a tool to flag AI‑generated posts LinkedIn announced a “Seems like AI slop” button on July 30, 2026...
- Anthropic’s Opus 4.6 Model Generates Explicit Sexual Content
Anthropic’s Opus 4.6 model can be coaxed into generating explicit sexual content despite official safeguards, highlighting a safety gap in legacy models.
Share your thoughts
Reactions, corrections, or insights — all welcome.
