Inclusion Health Screen a patient

Partnership brief

The tool works. The evaluation does not. That is what we are asking for.

For hospitals, trial sites, academic groups and clinical research units.

Document
Partnership brief, version 1.0
Audience
Principal investigators, research directors, trial site leads, CRU managers
The ask
Methodological review, advisory input, or a formal validation study
Cost to your institution
None. No licence, no fee, no purchase
Patient data required
None. See section 8
Reading time
About ten minutes

The ask, up front

We built a clinical trial eligibility screener that reads every recruiting study on ClinicalTrials.gov and tells one person which studies have already ruled them out. It runs entirely in a browser and collects nothing.

It has never been clinically validated. We can measure it against our own labels, and we have. What we cannot do alone is find out whether it helps a patient find a trial, or helps the coordinator who screens them afterwards.

That needs someone who has actually done the screening. The smallest useful version of this is one hour of a clinician's time reviewing the method. The largest is a co-authored validation study. Both are described in section 6.

We are not asking for money, for an endorsement, or for access to your patients. Section 10 lists what we are explicitly not asking for, because that question comes up first and deserves a direct answer.

What the tool is

A person types a condition. The tool retrieves every recruiting study for that condition from the ClinicalTrials.gov v2 API, reads the published eligibility criteria of each, and sorts the studies into those that have already excluded this person and those that have not. A solver then asks the smallest number of follow-up questions that resolves the most remaining studies.

Its distinguishing property is what it refuses to do. A deterministic rule compiler produces every verdict; the machine learning components are only permitted to group and surface criteria the compiler could not read. Criteria that delegate to investigator judgement are never machine-decided. Below a fixed confidence threshold the system abstains and shows the criterion verbatim instead of answering it.

Every search reports what fraction of criteria it managed to compile, so a reader can see how much of the text the tool did not read.

The full method note describes the architecture, the training procedure, the evaluation and the failure modes in detail.

Why this is your problem too

Two well-known frictions sit on either side of the same gap.

On the patient side

Eligibility criteria are public but unreadable. They are written for coordinators, published as unformatted prose, and a single record routinely contains dozens of separate conditions. People who might qualify never get as far as contacting a site, because the first step is reading something not written for them.

On your side

Recruitment is the constraint on most trials, and screening consumes coordinator time whether or not the person turns out to be eligible. Every enquiry that was never going to pass a published exclusion is time that produced nothing.

A screener that removes clearly ineligible enquiries before they reach a coordinator, while being conservative enough never to discourage someone who might qualify, would help both sides. Whether this tool actually does that is an empirical question that nobody has answered. That is precisely the gap we are asking you to help close.

What has been measured

Current internal measurements
QuantityValueAgainst
Compiler precision100.0%46 hand-labelled criteria
Conformal coverage92.9%same labels, alpha 0.10
Expected calibration error0.0008held-out weak labels
Training corpus32,628criteria, weakly labelled
Patient fields transmitted0by construction

The engineering is real and the numbers are honest. They are also, on their own, insufficient, for the reason in the next section.

Why those numbers are not enough

The 46 criteria were labelled by one rater, and that rater is one of us. Someone grading their own system, on a sample small enough that a single error would move the headline figure by more than two points.

That measurement does one useful job, which is catching a regression between builds. It cannot answer any of the questions a reviewer, a sponsor, or an ethics committee would ask:

  • Does it perform this way on criteria that a clinician, not the author, considers representative?
  • Do independent raters agree with each other, and where they disagree, does the tool's abstention behaviour match the disagreement?
  • Does it fail differently across conditions, phases, or sponsor types?
  • When it is wrong, is it wrong in a way that is merely unhelpful, or in a way that could send someone down a false path?

We would rather state this plainly than have a reviewer discover it. A tool that reports a confidence interval it cannot support is exactly the failure mode this project was built to avoid, and it would be incoherent to commit that error while describing the project.

Three ways to help

These are ordered by cost to you. The first is genuinely small, and it is the most common starting point.

Tier one

Read the method and tell us where it is wrong

One conversation. You use the tool for a few minutes on a condition you know well, look at what it refused to decide, and tell us what a clinician would consider dangerous, useless, or missing. The abstentions are the interesting part.

Cost: about one hour. No paperwork, no commitment, no name used anywhere.

Tier two

Advise on the evaluation design

Help specify what a defensible validation looks like: how criteria should be sampled, how many raters, what the adjudication protocol is, what the primary endpoint should be, and what result would count as a failure.

Cost: a few hours across several weeks. Acknowledged or co-authored, your choice.

Tier three

Run the validation study with us

A formal, pre-registered evaluation, described in the next section, with results published whichever way they come out.

Cost: depends on design. We do the engineering, the data collection tooling and the analysis.

The study we propose

This is a starting point for discussion, not a protocol. The design is exactly the thing we want a clinical partner to change.

Primary: criterion-level agreement

Design
Retrospective, criterion-level. No patients, no recruitment, no intervention.
Sample
A stratified sample of published eligibility criteria drawn across conditions, phases and sponsor types, sized for a usable confidence interval rather than for convenience.
Reference standard
Independent labelling by multiple clinicians blinded to the tool's output, with a pre-specified adjudication procedure for disagreements.
Primary endpoint
Agreement between the compiler's verdicts and the adjudicated human label, with abstentions treated as a distinct category rather than as errors or successes.
Secondary
Whether abstention concentrates where raters disagree, which is the behaviour the architecture predicts and the claim most worth falsifying.
Pre-registration
Analysis plan registered before labelling begins. Negative results published.

Optional second phase: usability

If the first phase holds up, a small prospective usability study: do people who used the tool arrive at a coordinator better prepared, and does that measurably change screening time or screen-failure rate? This phase needs your site in a way the first does not.

Data, governance and IRB

The first phase is designed to be as close to exempt as a study can be, because it uses no patient data at all.

Patient data
None. The primary study labels publicly published registry text.
PHI
None handled, stored or transmitted at any point.
Tool architecture
Runs client-side. Age, sex, country and every answer stay in the browser tab. Only the condition name reaches the public registry API.
Storage
There is no patient database. Nothing to breach, subpoena or de-identify.
Enforcement
A content security policy served with every page restricts outbound connections at the browser level, so the claim does not rest on our word.
Study data
Rater labels only. Held wherever your governance requires, including entirely on your infrastructure if that is simpler.

We expect your IRB or ethics committee to have the final say on classification and will follow whatever they determine. If a second phase involving patients proceeds, it would be designed under your institution's governance from the start rather than retrofitted.

What your institution gets

Authorship
Co-authorship on any resulting publication, on terms you set. If you would rather be acknowledged than authored, that is fine.
Control of the record
A negative result gets published the same as a positive one. That is a condition of the collaboration, not a concession.
The engineering
We build the labelling tooling, run the analysis, and do the writing. The work we are asking for is judgement, not labour.
The tool itself
Free, now and permanently, for patients and for your coordinators. There is no version of this where a validation partner is later asked to pay.
Direction
Real influence over what gets built next. The roadmap is currently set by two students' guesses about what a coordinator needs.

What this is not

Not a sales conversation
Nothing is for sale. There is no licence, no pilot fee and no procurement process.
Not an endorsement request
No name, logo or institutional affiliation appears anywhere without a written agreement, and none appears now.
Not a medical device
It gives no diagnosis, no advice and no probability of enrolment. A study that is not ruled out is not a study a person is eligible for.
Not a chatbot
There is no prompt in the product and no call to a language model API. The system is deterministic and versioned: the same input gives the same output, and a result stays reproducible.
Not asking for patients
The primary study needs no patient contact of any kind.

Honest risk register

The things most likely to go wrong, stated before you find them.

Known risks and current mitigations
RiskMitigation
The tool performs worse under independent labelling than internal numbers suggest That is the expected outcome of a real evaluation and the reason to run one. Results publish either way.
A person misreads "not ruled out" as "eligible" Stated on the landing page, in a blocking first-run agreement, in the Terms, and beside every result. Whether it works is itself worth measuring.
Registry records are stale or wrong Stated plainly to users, who are told to confirm with the site. The tool does not correct registry content.
The project is two people and could stall The engine, training pipeline and evaluation are readable and reproducible. Nothing depends on a hosted service that can disappear.
Lexicon coverage is thin for rare diseases Coverage is reported on every search rather than hidden. Improving it is a concrete thing a partner can direct.

Who is asking

We are Akshobya Rao and Abhinav Byju, juniors at Neuqua Valley High School in Naperville, Illinois, and the two co-founders of Inclusion Health. We built the whole thing ourselves: the rule compiler, the training pipeline, the calibration and the interface.

We are telling you this at the start rather than letting you find it out later. The work should be judged on whether the method holds, and the fastest way to check that is to spend two minutes with the tool on a condition you know well and look at what it refused to decide.

How to start

The fastest way to judge this is not to read more. Open the tool, screen a condition you know, and open a study to see what it refused to answer. The abstentions say more about the system than the verdicts do.

If it looks worth an hour, reply and say which of the three tiers in section 6 fits. If it looks wrong, saying so plainly is more useful than silence, and we would rather hear it now.

Or write to hello@inclusionhealth.xyz. There is no form, no calendar link and no follow-up sequence.