VMTech
Discuss a project →

Inherent says Faraday AI agent surpassed larger models in paper replication

Inherent says Faraday AI agent surpassed larger models in paper replication

London AI lab Inherent, founded by Google DeepMind alumni, says its Faraday agent has outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at independently reproducing the findings of published scientific papers. The startup says Faraday performed the task using Qwen 3.6, a 27-billion-parameter model, rather than a frontier-scale system.

The claim comes weeks after Inherent emerged from stealth with a $50 million seed round. The company is pursuing a broader objective than result verification: it wants to build an AI scientist agent capable of helping discover new scientific knowledge across fields.

Paper replication as a test of scientific work

Reproducing a paper’s findings without being given the answer in advance is a useful exercise for assessing an agent’s ability to conduct research. Inherent cofounder and chief scientist Edward Hughes compared it with a common starting point for PhD students, who often learn by replicating existing work before pursuing original questions.

For Inherent, the reported benchmark result was less significant than the method used to reach it. The company set a bar beyond accuracy: Faraday was expected to show what it calls research taste, including judgment about which experiments are worth running and how to design them.

Reinforcement learning and external tools

Inherent says it relies on reinforcement learning to develop that judgment. This approach rewards an AI system for desirable outcomes rather than supplying a fixed set of rules. The company is betting that reward-based training will generalise more effectively to its longer-term goal of scientific agents operating across multiple disciplines.

Faraday was not designed to rebuild every capability internally. For coding work, Inherent had the agent use OpenAI’s GPT-5.5 Codex, mirroring the way researchers use available software tools instead of developing each component themselves. The company frames Faraday as a teammate that investigates questions, runs experiments and returns with results for discussion, rather than an agent optimised merely to validate a user’s preference.

London hiring plans

Inherent has about a dozen employees working in person from King’s Cross, London, and plans to expand to roughly 20 to 25 people by the end of the year. Hughes has also argued against UK “garden leave” restrictions, which can prevent departing employees from joining or founding rival companies for months.

The practical implication for businesses evaluating research agents is to test their experimental choices, reproducibility and use of specialist tools alongside output accuracy. A smaller underlying model may still support a capable workflow when the agent’s training and tool orchestration fit the task.

#artificialintelligence#aiagents#reinforcementlearning#research

How to interpret Inherent’s Faraday research result

The Inherent Faraday AI agent is presented as more than a system for generating answers. Its reported paper-replication workflow combines experimental planning, reinforcement learning and specialist tools. The important question is therefore not only whether Faraday reached the expected result, but whether its process was reproducible, well designed and appropriately supervised.

What the Faraday agent is designed to do

Inherent describes Faraday as an AI scientist and research teammate that can investigate a question, select experiments, use external tools and return results for discussion. In the reported replication task, the agent used Qwen 3.6, a 27-billion-parameter model, while relying on GPT-5.5 Codex for coding work rather than recreating every capability internally.

  • Faraday is intended to plan and conduct research tasks, not merely summarise papers.
  • The workflow combines an underlying model with training, tools and experimental choices.
  • Human review remains relevant when assessing methods, evidence and conclusions.

What the reported comparison means

Inherent says Faraday surpassed larger Anthropic and OpenAI models when reproducing findings from published papers. That is a company-reported benchmark result, not proof that the Faraday AI agent is universally more capable. Performance can depend on the selected papers, evaluation criteria, available tools and the amount of guidance given to each system.

  • A replication result should be interpreted within its stated test conditions.
  • Model size alone does not describe the capability of an agent-based workflow.
  • Independent repetition would make the comparison easier to assess.

How to evaluate an AI research teammate

For an Inherent-style AI teammate, research replication should be judged at the workflow level. Reviewers should examine whether the agent chose informative experiments, documented its steps, handled failed attempts and produced results that another researcher could reproduce.

  • Check whether experimental choices directly test the paper’s central claims.
  • Record prompts, tools, code, data and configuration needed to repeat the work.
  • Compare successful outputs with failed runs and unsupported conclusions.
  • Separate accurate reproduction from genuinely new scientific discovery.

Who is behind Inherent AI

Inherent is a London startup founded by Google DeepMind alumni. Its broader ambition is to develop an AI scientist capable of supporting research across disciplines. Faraday’s paper-replication exercise should be read as an early test of that direction rather than evidence that autonomous scientific discovery has already been achieved.

  • The company is building Faraday around research workflows and tool use.
  • Paper replication tests only part of the wider scientific process.
  • Claims about broader scientific capability require broader evidence.

Frequently asked questions

What is Inherent Faraday?

Faraday is an AI research agent developed by Inherent. The company presents it as a research teammate that can investigate questions, plan experiments, use external tools and report results.

Is Faraday already an autonomous AI scientist?

Inherent’s long-term objective is an AI scientist, but the reported result concerns replication of published research. It does not by itself establish autonomous discovery across scientific fields.

Why is the 27-billion-parameter model notable?

Inherent says Faraday used Qwen 3.6 with 27 billion parameters while outperforming larger models in its replication test. The claim suggests that training and tool orchestration may matter alongside model scale, but it remains specific to the reported evaluation.

Why does Inherent call Faraday a research teammate?

The teammate framing emphasises investigation, experiment design and discussion of results rather than simple answer generation. It also implies a workflow in which researchers review the agent’s methods and conclusions.

Open analytics
On the site 128 views
min read 3 22.08.2026
On Instagram 4 views
On Instagram 1 reach
Instagram

Inherent says Faraday AI agent surpassed larger models in paper replication

Open the post on Instagram ↗