Claude Code × Medical Application
[Claude Code] Systematic Review × AI Subagents Part 4: Discussion and How to Reproduce

1. Introduction
This is the final part of the series. In Parts 1–3, I rebuilt the original paper TrialMind’s three tasks (search, screening, and data extraction) with Claude Code and compared each with the original paper. Part 4 brings them together to answer this trial’s question, and shows how readers can run the same flow from the repository themselves.
1-1. Purpose of This Part
The question set in Part 0 was: “If the original paper’s method is rebuilt with Claude Code and measured with the same metrics, does it reach the same level as the original paper?” This part covers three things.
| Content | |
|---|---|
| Discussion | What can and cannot be said when the three tasks are put side by side with the original paper. Where human judgment matters, made visible by measuring without it |
| Inventory of the build | The subagents, skills, hooks, and MCP used in the series, and where the agents did not follow the procedure |
| How to reproduce | Why only a human launches the evaluation skill /eval, the setup for running on a subscription, and how to clone the repository and run the same flow |
In this part, “reproduce” is used only to mean readers running the same flow from the repository. As stated in Part 0, this trial is not a replication of the original paper, and I do not write that “the original paper was reproduced.”
1-2. Summary of the Three Tasks

| Task | Metric | Original paper | This trial (no human judgment) | Part |
|---|---|---|---|---|
| Search | Search Recall | 0.711–0.834 | 0.963 (26/27) | Part 1 |
| Screening | Recall@20 | 0.567 | 0.482 (mean of 3 reviews; pooled 14/27) | Part 2 |
| Screening | Recall@50 | 0.713 | 0.779 (mean of 3 reviews; pooled 22/27) | Part 2 |
| Extraction | Accuracy | 0.78 (patient characteristics 0.74) | 0.735 (75/102; 2 reviews, 7 pairs) | Part 3 |
The original paper’s values are from the arXiv HTML version: search Recall is the range across four topics, and the others are Immunotherapy values. Recall@50 and extraction (compared with the patient characteristics value of 0.74) are on par with the original paper; Recall@20 is lower. Search Recall is higher in number, but the original paper does not say how many candidates it narrowed to, so the two values cannot be compared (Part 1).
2. What Can and Cannot Be Said Against the Original Paper
What can and cannot be said for each task is in chapter 5 of each part. Taken together, what can be said comes down to one thing: the original paper’s three tasks could be rebuilt with Claude Code subagents, skills, and hooks, and measured with the same yardstick they produced values at a similar level. The original paper’s method worked in a similar way with a different model and a different toolset, at least on 3 reviews (2 for extraction).
What cannot be said is also shared across the three tasks.
- It cannot be said that performance matches the original paper. There are 3 reviews, 2 for extraction (the original paper used 100); the model is the current Claude Sonnet (the original paper used GPT-4 and Claude 3 Sonnet); and extraction had 1 scorer (the original paper had 3)
- It cannot be said that this is usable for real SR work. Search covers PubMed only, full text comes from PMC only, and the extraction items include no outcomes. As positioned in Part 0, this is a teaching tool and a tool for checking a method
- It is the result of one run. PubMed’s relevance order changes from call to call, and the same query returned results in a different order (Part 1). Running the same procedure will not necessarily give the same numbers
3. Where Human Judgment Matters

To compare with the original paper, this trial was measured on a flow with human judgment removed. Removing it is exactly what made it visible where human judgment has an effect.
| Step | What happened without a human | If a human steps in |
|---|---|---|
| Search | 26/27 with a query built without human input. The repository’s records also note that constraints a human added to the query at the approval step (block structure, specific terms) may have reduced coverage (not verified) | When adding a constraint, check by hit counts how much it narrows what can be found |
| Screening | Criterion I5 (comparison group) in the draft criteria remained, and single-arm trials each lost a point. Recall@50 for 31190844 was 0.429 (0.714 as a reference value with I5 removed) | A human decides the “questions a human should decide” in the draft, before judging |
| Extraction | What went into the input (affiliations, supplementary materials, figure images) set the ceiling on the values that could be extracted | A human decides the scope of the input to fit the items |
| Answers | For 37168849, the answers conflicted with the original review’s own criterion (CAR-T); 33495835 is in vitro, yet the answers had patient characteristics | Don’t fix the answers; report the conflicts separately |
The clearest case was I5 in screening. The skill that wrote the draft criteria left I5 as a question where “a human decides whether to apply it.” A question a human would have removed lowered the ranking as is in the flow without a human. It was not an agent error; a decision meant to be handed to a human never reached one.
Conversely, there is also a place where a human must not step in: moving the query, criteria, or rules after seeing the answers. In this trial, fixing things after seeing the answers would make the evaluation meaningless, so every step was recorded as it was. The rule throughout the series: human judgment goes in before the answers are seen.
4. How This Series Was Built
4-1. What Was Used

| Kind | Name | Step | Role |
|---|---|---|---|
| subagent | query-builder |
Search | Tries queries using only the PubMed connector’s two tools |
| subagent | screener |
Screening | 20 records at a time; 1 / 0 / −1 per criterion with verbatim quotes |
| subagent | extractor |
Extraction | Once per pair; a value and verbatim quotes per item |
| skill | /pico-to-criteria, screening-rules and extraction-rules, /prisma-record and /eval |
Throughout | Draft criteria, preloading rules, PRISMA counts, recording evaluations |
| hook | check_screen_output.py, check_extract_output.py, limit_reads.py |
Screening, extraction | Checking output and verbatim quotes, restricting readable files (PreToolUse) |
| hook | agent_gate.py, check_prisma.py, log_prompt.py |
Throughout | Concurrency limit, PRISMA sums, logging human instructions |
| MCP | PubMed connector (official) | Search | Tries partial queries. Bulk retrieval is done by a script |
The pattern is the same across the three tasks: hand work to a subagent one unit at a time, preload rules with a skill, check output with a hook right before it is written, and leave counting to scripts. What I left to the agents was only “read and write”; every number was copied from script output.
4-2. Where Things Did Not Go by the Procedure
| Part | What happened | How it was handled |
|---|---|---|
| Part 1 | The query-builder definition says “pick up terms from the abstracts of a preliminary search,” but it built queries without reading abstracts |
Kept the record unchanged and wrote it under what cannot be said |
| Part 2 | The rule is that only a human launches /eval, but the eval-3 evaluation was run directly as scripts by another session, on a human’s instruction |
Recorded that results, criteria, and queries were not changed |
| Part 3 | An agent definition added mid-session was not loaded, and the trial run stopped. The extractor’s report and its output disagreed on a count | Restarted rather than substituting. Treated the output, not the report, as correct |
| Throughout (a run before eval-3) | The main agent miscounted the running subagents and nearly launched a 7th | Claude Code’s built-in limit stopped it |
| Preparing for release | Rewriting history (git filter-repo) was stopped by auto mode’s safety check |
A human typed it in the terminal |
What a definition says and what actually happens are different things. I noticed these because the output was kept in files, checked with hooks and scripts, and anything unexpected was written on the spot as one line in docs/HARNESS.md.
5. Why Only a Human Launches /eval
5-1. Setup
The evaluation skill lives in .claude/skills/eval/SKILL.md, and its frontmatter sets disable-model-invocation: true. With this, Claude cannot call the skill on its own judgment; it runs only when a human types /eval.
name: eval
disable-model-invocation: true
argument-hint: <eval の名前(例:eval-4)> [--run eval-3] [--results <dir>]
(The argument hint reads “<name of the eval (e.g., eval-4)>”.) The procedure has these steps: 0 check prerequisites (no record with the same name, zero check warnings) → 1 write what will be measured first → 2 copy from script output (no hand counting) → 3 compare with the previous run and the original paper → 4 classify what was missed → 5 record and stop. Commits and tags wait for a human’s instruction.
5-2. Why
Evaluation is the step where the answers (the PMIDs of included studies and the extraction answers) are used for the first time. SKILL.md also says that the answers are used for the first time at this step, and that judgments and extracted values are not adjusted to match them.
If an agent can run the evaluation on its own judgment, it gets the chance to look at the answers midway and fix the query or criteria. If a human controls when the answers are seen, the procedure before seeing the answers and the record after seeing them stay separate. This locks in, as a setting, Part 0’s rule that “only a human launches a full evaluation run.”
However, as the table in 4-2 shows, the eval-3 evaluation itself did not go through this skill. Another session ran the same scripts directly on a human’s instruction. The setting only stops “Claude calling it on its own”; it does not stop someone running the same scripts on a human’s instruction. The last safeguard was keeping a record.
5-3. A Fresh Clone Alone Cannot Be Evaluated
On a fresh clone, the commands that step 0 of /eval uses to check prerequisites both stop with this line:
results/eval-3/ が無い。results/ は commit されない。CLAUDE.md の「データの流れ」の順に作る
(“results/eval-3/ does not exist. results/ is not committed. Build it in the order of ‘Data flow’ in CLAUDE.md.”) Judgments, candidates, and extraction scores are written to results/ and are not committed. The evaluation is something you run after running the flow yourself and building results/. There is not yet a record of a human typing /eval itself to confirm how it behaves.
6. Setup for Running on a Subscription
6-1. No API Key
This repository assumes Claude Code runs on a Pro or Max subscription. README’s “Requirements” includes this line in bold:
Do not set
ANTHROPIC_API_KEY(setting it switches to pay-as-you-go API billing)
According to the official documentation, if the environment variable ANTHROPIC_API_KEY is present, Claude Code uses that key ahead of the subscription. In interactive mode, you are asked once whether to use it, and the answer is remembered. In non-interactive mode (-p), it is used without asking. In a flow that launches subagents tens to a hundred times, approving it once by accident means the whole run is billed per use. You can check which one is in use with /status. My instruction sheets also listed “stop if it is set” as a stop condition, and before the full runs in Parts 2 and 3 I checked and recorded that it was “not set” (checking only whether it exists, without looking at the value).
6-2. A Double Concurrency Limit
Subscriptions have usage limits. Launching many subagents at once reaches the limit sooner and stops midway. .claude/settings.json limits concurrent subagents to 6.
"env": {
"CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS": "6",
"CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH": "1"
}
On top of this, the SubagentStart hook agent_gate.py counts running subagents with marker files and, if 6 are running, stops the 7th from launching with exit 2. There are two reasons for setting the limit twice, in Claude Code’s settings and in a project hook: the hook can stop it with a reason (“the limit in CLAUDE.md is 6”), and the hook also counts cases the built-in limit does not (resuming a finished subagent). In practice, when the main agent miscounted and nearly launched a 7th, the built-in limit stopped it first (4-2). SPAWN_DEPTH=1 keeps subagents from launching further subagents.
6-3. Rough Amounts
The amounts on record are as follows (subagent-side tokens; the main agent’s share is not included).
| Step | Per run | Runs |
|---|---|---|
| Screening (screener, 20 records) | Trial run: about 93 s, about 36k tokens | 95 batches across 3 reviews |
| Extraction (extractor, 1 pair) | 16.7–30.3 s, 15k–31k tokens | About 177k tokens for 7 pairs |
Splitting work into subagents does not stop the total from growing with the number of records. When running it yourself, check the amount with a one-batch trial run before moving to the full set.
7. Running the Same Flow from the Repository

7-1. Setup
You need Claude Code (Pro or Max), Python 3, and the official PubMed connector. An NCBI API key and Node.js are optional.
git clone https://github.com/HerzLeben/pubmed-slr-screening.git
cd pubmed-slr-screening
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt -r requirements-dev.txt
.venv/bin/python -m pytest -q
.venv/bin/ruff check .
The tests do not use the network. Install the PubMed connector from inside Claude Code.
/plugin marketplace add anthropics/life-sciences
/plugin install pubmed@life-sciences
After installing, restart Claude Code and check the connection with /mcp. During the series, the connector’s tools were sometimes not visible in the VS Code extension’s panel but were visible in an interactive terminal session. If you can’t see them, try claude in the terminal.
7-2. Data
The evaluation data, TrialReviewBench, is not in the repository. Fetch it into bench/raw/ with the three curl commands in README, then format it with build_bench.py and build_extraction.py (the formatted files are also in the repository; these two commands rebuild them).
7-3. Running the Flow
From here on, you proceed by instructing Claude Code, not by typing commands. The “Data flow” section of CLAUDE.md lists which script or subagent reads and writes what at each step, so you give instructions in that order.
| Step | What to ask Claude Code | Where it is written |
|---|---|---|
| Candidates | Fetch all hits from reviews/<rid>/eval-3/search.json (fetch_pubmed.py --from-search ... --all-hits --out-dir results/eval-3; for 37168849, also --add-pmid 28864289) |
results/eval-3/<rid>/ |
| Screening | Build batches with make_batches.py --review-pmid <rid> --run eval-3, try screener on one batch → all batches |
results/eval-3/screen/a/ |
| Extraction | fetch_pmc.py --from-extraction 33746596 37168849 → fulltext_to_text.py → make_extraction_jobs.py → try extractor on one pair → all pairs |
results/fulltext/, results/extraction/ |
| Scoring | Build the screen with build_report.py --run eval-3; a human scores and saves |
results/extraction/human/ |
| Evaluation | A human types /eval <name> --run eval-3 |
docs/eval/<name>.md |
The query and draft criteria are committed in reviews/<rid>/eval-3/, so the table above uses them and runs everything from search onward. If you also want to rebuild the query yourself, start by giving the PICO to query-builder. Abstracts, full text, and judgment results are written to results/ and are not committed.
7-4. You Will Not Get the Same Numbers
Running the same procedure will not necessarily give the same numbers as Parts 1–3.
- Bibliographic records such as abstracts may change depending on when they are fetched (the candidate order is fixed by the committed search.json; rebuilding from the query changes the order too)
- Judgments and extraction are model output and vary from run to run
- Extraction is scored by a human
At the time of preparing for release, I confirmed on a fresh clone that setup (venv → requirements → pytest → ruff) and rebuilding bench/ work. However, there is no record of running Claude Code end to end from search to extraction again. If the numbers differ, that in itself is material for checking the method.
8. Using It for Teaching and Verification
The positioning has not changed since Part 0: it is a teaching tool for learning SR methods and a tool for checking a paper’s method yourself. It is not a tool distributed for real work. Some ways to use it (none of these were tried in the series):
- Compare the draft criteria with the human-edited version: with
git diff snap/03-criteria-draft snap/04-criteria-approved -- reviews/, read what a human changed at approval and where questions like I5 were - Swap in another review: pass a different PMID from TrialReviewBench’s 100 reviews to
build_bench.py. Check first whether extraction answers exist for it - Change the model: change
modelin the agent definitions and see how Recall@k moves on the same candidates - Read the records:
docs/HARNESS.mdanddocs/prompts/log.mdare real examples of where things get stuck when you leave a procedure to an agent
9. Summary
The answer to this trial’s question is: when the original paper’s three tasks are rebuilt with Claude Code, they produce values at a level close to the original paper on 3 reviews (2 for extraction). Recall@50 of 0.779 and extraction Accuracy of 0.735 (original paper’s patient characteristics: 0.74) are on par with the original paper; Recall@20 of 0.482 is lower. Search Recall of 0.963 was measured under different conditions and cannot be compared. Still, the number of reviews, number of candidates, model, and number of scorers differ, so it cannot be said that performance matches the original paper, nor that it is usable for real work.
Measuring without human judgment made it visible where human judgment matters: questions left in the draft criteria, the scope of the input, and conflicts in the answers. And human judgment goes in before the answers are seen. That is why only a human launches the evaluation skill /eval.
In terms of structure, I handed work to subagents one unit at a time, preloaded rules with skills, checked output with hooks, and left counting to scripts. Writing something in a definition does not mean an agent will do it. So keep records, and build the mechanisms that check them first. How much you can delegate is decided by how verifiable you made it in advance.
Next Steps
This is the end of the series. The code is on GitHub (MIT), and the results are on the sample screen. If you reread from Part 0, putting each part’s “what cannot be said” side by side gives the clearest picture of this trial’s scope.
Sources
- Wang Z, Cao L, Danek B, Jin Q, Lu Z, Sun J. Accelerating clinical evidence synthesis with large language models. npj Digit Med 2025;8:509, doi:10.1038/s41746-025-01840-7 (arXiv:2406.17755)
- TrialReviewBench (Hugging Face: zifeng-ai/TrialReviewBench, Apache-2.0)
- Anthropic, Claude Code Docs (subagents / skills / hooks / settings)
The code is available on GitHub: github.com/HerzLeben/pubmed-slr-screening
