The Claude enzyme discovery is a result Anthropic announced on September 23, 2026: roughly 950 Claude agents searched public genomic databases for 21 hours and identified a previously uncharacterized enzyme system in bacteriophages. The system, named ART (array-associated reverse transcriptases), pairs a reverse transcriptase gene with a long array of evenly spaced DNA repeats that looks structurally like a CRISPR array. Nobody yet knows what it does.
That last sentence is the one to hold on to. Headlines called it "the next CRISPR." What the agents found is a pattern — a genetic architecture that resembles CRISPR's layout — not a working gene editor, and not a function. The pre-print has not been peer reviewed.
It is still a significant result, for a reason that has little to do with biology hype: the humans supplied the initial prompt and did the lab work, and the agents did the rest. This article covers what ART is, how the search was run, what it likely cost, what has and has not been validated, and what it means for agent-driven research.
Key Takeaways
- About 950 Claude agents ran for 21 hours and used roughly 210 million tokens to search public sequence databases.
- They gathered more than 200,000 reverse transcriptases, flagged 3,500 candidate systems and wrote detailed reports on the 20 most compelling.
- ART has three parts: a reverse transcriptase, a partner gene of unknown function, and a CRISPR-like array of evenly spaced repeats.
- Early lab work shows the ART array is expressed as a set of distinct short RNAs, but the system's biological function is unknown.
- CRISPR pioneer Feng Zhang called the finding "genuinely intriguing" and said it merits further investigation; the pre-print is not yet peer reviewed.
What did Claude actually discover?
Claude agents discovered ART, a previously uncharacterized enzyme system found mainly in bacteriophages — viruses that infect bacteria. It has three components: a reverse transcriptase enzyme, a neighbouring accessory gene of unknown function, and a long array of evenly spaced DNA repeat sequences. The layout resembles a CRISPR array.
To see why that is interesting, it helps to know what each piece normally does.
Reverse transcriptases are enzymes that copy RNA into DNA, the reverse of the usual direction. They are the machinery behind retroviruses, and they are also the key component in newer gene-editing techniques such as prime editing.
CRISPR arrays are stretches of DNA in bacteria made of short, identical repeats separated by unique "spacer" sequences. Each spacer is a fragment of a virus the cell has met before — a molecular memory that lets the cell recognize and cut that virus next time. The discovery that this system could be reprogrammed is what produced modern gene editing.
ART puts those two things next to each other: a reverse transcriptase sitting beside a repeat array. According to Anthropic's announcement, the agents flagged the family when they noticed "a tandem repeat array … that's a CRISPR-like … repeat array?!" next to an unusual reverse transcriptase.
Two details complicate the easy comparison.
First, ART lives mainly in phages, not bacteria. CRISPR is a bacterial defence against phages. Finding a CRISPR-like architecture on the virus side is the reverse of the familiar arrangement, and nobody knows yet whether it is a weapon, a counter-defence or something unrelated.
Second, the resemblance is architectural. The array looks like a CRISPR array. That does not mean it works like one.
How did 950 agents search 200,000 enzymes?
The agents worked as a large parallel search. They collected more than 200,000 reverse transcriptase sequences from public databases, examined the genes surrounding each one, identified 3,500 new candidate systems, and then analyzed the 20 most compelling in depth, producing human-readable reports. The run used about 210 million tokens over 21 hours.
The funnel looks like this:
| Stage | Count | What happened |
|---|---|---|
| Sequences gathered | 200,000+ | Reverse transcriptases pulled from public sequence databases |
| Candidate systems | 3,500 | Novel gene neighbourhoods flagged as potentially new systems |
| Deep-dive reports | 20 | The most compelling candidates written up for human review |
| Named system | 1 | ART, taken forward to laboratory work |
The method is a familiar one in computational biology, sometimes called "guilt by association." Genes that work together tend to sit next to each other in microbial genomes. If you find an enzyme that always appears beside the same unusual neighbour, the pair probably forms a functional system. This is broadly how many CRISPR-associated systems were found in the first place.
What is new is who did the looking. A human bioinformatician can run this kind of neighbourhood analysis. What a human cannot do is read 3,500 candidate neighbourhoods with full attention, form a hypothesis about each, check it against the literature and decide which 20 deserve a report — in under a day.
The core pattern the agents were hunting for is simple enough to sketch. Here is a toy detector for evenly spaced repeats in a DNA string:
from collections import defaultdict
def find_repeat_arrays(seq, k=8, min_copies=4, tolerance=2):
"""Find k-mers that recur at near-regular intervals (a repeat array)."""
positions = defaultdict(list)
for i in range(len(seq) - k + 1):
positions[seq[i:i + k]].append(i)
hits = []
for kmer, pos in positions.items():
if len(pos) < min_copies:
continue
gaps = [b - a for a, b in zip(pos, pos[1:])]
regular = max(gaps) - min(gaps) <= tolerance
if regular and min(gaps) > k: # repeats separated by spacers
hits.append((kmer, pos[0], len(pos), gaps[0]))
return hits
repeat = "GTTTCAGA"
spacers = ["ACGTACGTTAGC", "TTGACCATGCAA", "CGATCGGATTAC", "AAGCTTGCATGC"]
seq = "ATGC" * 5 + "".join(repeat + s for s in spacers) + repeat + "TTAA" * 5
for kmer, start, copies, spacing in find_repeat_arrays(seq):
print(f"{kmer} starts at {start}: {copies} copies, every {spacing} bp")
# GTTTCAGA starts at 20: 5 copies, every 20 bp
Real tools handle mismatches, reverse complements and genome-scale data, and this sketch does none of that. But it shows the shape of the signal: identical units, regular spacing, unique sequence in between. The hard part of the discovery was not detecting repeats. It was deciding which of thousands of odd neighbourhoods were biologically interesting.
What did the search cost?
Anthropic has not published a cost, but the token count gives a bound. At Claude Opus 5.5 list prices of $4 per million input tokens and $20 per million output tokens, 210 million tokens cost between about $840 (if every token were input) and $4,200 (if every token were output). The real figure sits somewhere in between.
This is our own arithmetic, and it carries assumptions: Anthropic did not say which model the agents ran on, and it ignores caching discounts and any internal pricing. Treat it as an order of magnitude.
TOKENS = 210_000_000
PRICE_IN, PRICE_OUT = 4.00, 20.00 # USD per 1M tokens, Claude Opus 5.5 list
low = TOKENS * PRICE_IN / 1_000_000
high = TOKENS * PRICE_OUT / 1_000_000
print(f"${low:,.0f} to ${high:,.0f}") # $840 to $4,200
agent_hours = 950 * 21
print(f"{agent_hours:,} agent-hours") # 19,950 agent-hours
The order of magnitude is the point. A search that screened 200,000 enzymes and surfaced a candidate system drew a comment from one of the inventors of CRISPR gene editing — for what is plausibly a four-figure compute bill and under a day of wall-clock time.
Compare the other side of the ledger. If all 950 agents ran for the full 21 hours, that is 19,950 agent-hours of analysis. A single researcher working 40-hour weeks would need roughly nine and a half years to log that many hours.
That comparison is imperfect — an agent-hour is not a scientist-hour — but it explains why this matters more than the specific enzyme. The bottleneck in this kind of discovery has moved. Searching is now cheap. Validation is the expensive part, and it still happens at the speed of a wet lab.
Has the Claude enzyme discovery been validated?
Partially. Human scientists at Anthropic's life sciences lab ran initial experiments and found that the ART array is expressed as a set of distinct short RNAs, which suggests the array is biologically active rather than genomic debris. The function of the system remains unknown, and the pre-print has not been peer reviewed.
Here is an honest scorecard.
| Claim | Status |
|---|---|
| The ART architecture exists in phage genomes | Supported by sequence analysis |
| The array is transcribed into short RNAs | Supported by initial lab experiments |
| ART has a CRISPR-like function | Unknown |
| ART can be used for gene editing | Not shown |
| Independent replication | Not yet |
| Peer review | Not yet |
The lab work was done by people. Anthropic's lab is in the Bay Area and operates at biosafety levels BSL-1 and BSL-2. The company is explicit that humans directed "the initial prompt and the lab work," while the agents ran the search autonomously.
Outside reaction has been measured. Feng Zhang, the MIT and Broad Institute professor who helped develop CRISPR genome editing, reviewed the work and said: "This is an exciting example of how AI agents can contribute to biological discovery." He added that "the identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation."
Read that carefully. "Merits further investigation" is a scientist's way of saying the observation is real and the interpretation is open. As Tech-ish put it in its headline, Claude found the system "but nobody knows what it does yet."
Interesting Engineering's coverage makes the same distinction. Plenty of intriguing genomic architectures turn out to be minor curiosities. Some turn out to be CRISPR. You cannot tell which from the sequence alone.
Why does this matter for AI agents?
It matters because it is a concrete example of agents producing a novel scientific observation with minimal human steering, at a scale no human team could match in the same time. The specific enzyme may or may not prove important. The demonstration that a swarm of agents can run an open-ended search and return something experts find worth testing is the durable result.
We have now seen this in three different fields within a few months. Models produced new results on open mathematics problems, covered in AI solved Erdős problems: what actually happened and in OpenAI Astra's math proofs. Now there is a biology result. The same week, Claude was reported to have completed a nine-loop physics calculation one step beyond the previous human record.
The common structure across all three is worth naming, because it tells you what kind of problem agents are currently good at:
- A huge search space that is tedious rather than conceptually impossible.
- A cheap way to check candidates computationally before anything expensive happens.
- An expensive final verification step — a proof check, an experiment — that a human or a lab still performs.
Biology fits that pattern well, and it has a constraint the other two fields do not. A proof can be verified in a proof assistant. An enzyme has to be expressed, purified and tested. The result is a widening gap between how fast hypotheses can be generated and how fast they can be checked.
That gap is the unique risk here. If agents can produce 3,500 candidates in a day and a lab can test a handful in a month, the scarce resource becomes experimental capacity and the judgement to choose what to test. Labs that build faster validation will get more from these tools than labs that simply run bigger searches.
There is a safety dimension too. The same capability that finds a useful enzyme system can, in principle, surface harmful biology. Anthropic gates advanced biology capability behind a Life Sciences Verification Program for its newest models. If you build agent systems yourself, our guide to AI agent frameworks covers how orchestration at this scale is structured.
Frequently asked questions
What enzyme did Claude discover? Claude agents discovered ART, short for array-associated reverse transcriptases. It is a previously uncharacterized system found mainly in bacteriophages, made up of a reverse transcriptase, a partner gene and a CRISPR-like array of evenly spaced DNA repeats.
Is ART the next CRISPR? Not yet, and possibly not at all. ART resembles CRISPR in its genetic layout, but its biological function is unknown and no gene-editing use has been demonstrated. The work is a pre-print that has not been peer reviewed.
How many Claude agents were used in the discovery? About 950 Claude agents ran for 21 hours, using roughly 210 million tokens. They gathered more than 200,000 reverse transcriptases, identified 3,500 candidate systems and produced detailed reports on the 20 most promising.
Did Claude do the lab experiments? No. All laboratory work was carried out by human scientists at Anthropic's Bay Area lab, which operates at biosafety levels BSL-1 and BSL-2. The agents performed the computational search and analysis.
What did Feng Zhang say about the Claude discovery? Feng Zhang called it "an exciting example of how AI agents can contribute to biological discovery." He said the identification of RNA-repeat arrays associated with reverse transcriptases "is genuinely intriguing and merits further investigation."
What is a reverse transcriptase? A reverse transcriptase is an enzyme that copies RNA into DNA, the opposite of the usual direction of genetic information. These enzymes are used by retroviruses and are a core component of gene-editing techniques such as prime editing.
The verdict
The Claude enzyme discovery is real, interesting and unfinished. ART is a genuine, previously undescribed genetic architecture, the early lab data suggests it is active, and a leading figure in the field thinks it deserves a closer look. It is not a gene editor, and "the next CRISPR" is a headline rather than a finding.
The part that will age well is the method. A sub-day search across 200,000 enzymes for a plausibly four-figure compute cost means the limiting factor in this kind of science is no longer finding candidates. It is testing them. Whoever builds faster experimental loops will capture most of the value these agents create.
If you want the pattern behind this result rather than the biology, start with our piece on how AI solved Erdős problems — the structure is the same, and so are the caveats.
The agents read 200,000 sequences in a day. Working out what one of them means is going to take people a lot longer.