The Engine — distill the record, hunt the ghosts

28 million abstracts. One question. One morning.

PDAOAI folds the ~28M abstracts indexed by the National Library of Medicine into a single searchable space. Building the index takes months and never stops; putting a question to it takes a morning of compute. The first open question it was asked — why did rapalogs fail in liver cancer for twenty years? — came back with a candidate patient group the trials never defined. No wet lab, no new patients. Manuscript submitted for publication.

Built on Qdrant
The Corpus
28M

PubMed abstracts — the published record of biomedicine, in one searchable space.

The ~28M abstract-carrying citations indexed by the U.S. National Library of Medicine. Ingested, embedded in a Qdrant vector layer, folded to make convergence visible, and refreshed on a rolling basis as new science is indexed. This is the substrate PDAOAI builds on.

Why single-query AI fails here

Ask 28 million abstracts one question, and the loudest disease answers it.

Which cancers respond to rapalogs? Ask the whole corpus at once and the answer comes back weighted by how much each disease publishes.

HR+ breast cancer 588,114
Mantle cell lymphoma 7,979
74× more literature — not 74× more clinical signal. A single query cannot tell the two apart. So Partitioned Literature Aggregation splits the corpus into 23 disease-specific bodies of evidence, asks each one the identical question, and only then aggregates. The quiet diseases get heard.

4,453,833 rapalog-relevant abstracts, partitioned into 23 disease corpora. It is not a search problem. It is a signal problem.

Run 01 · YAP1 · June 2026

No paper called YAP1 the master driver. The corpus did.

Taxane resistance has been linked to more than a hundred genes across drug efflux, EMT, mTOR signalling, metabolism and microtubules. No individual paper ties the picture together, and the field does not have a named master driver.

What the engine measured
Sprawl score 0.92Every candidate scored on how far its literature reaches across independent mechanisms. YAP1 came back at 2.5–4× every other gene in the set.
Five mechanistic domainsDrug efflux, EMT, mTOR signalling, metabolism and mechanotransduction — a footprint no single paper claims together.
How it was checked
Held out, then testedPatient-outcome data was excluded from the ranking and used afterwards as a check — a four-gene effector panel against public taxane-treated ovarian cohorts, HR 1.17–1.45, univariate and unadjusted.
Filed, then disclosedProvisional filed 10 June 2026. Presented publicly at Vector Space Day on 11 June. One day apart, in that order.

No paper claims all five. The structure of the corpus does. Found from the literature alone, then confirmed against outcomes it had never seen.

Case Study — mTOR in Hepatocellular Carcinoma

A patient group two decades of trials never found. One morning of compute surfaced it.

That is the question the engine was pointed at — the first with no known answer. Twenty years of rapalog trials in liver cancer had failed, and nobody could say whether the drug class was wrong or the patients were. The partitioned run above is how it was asked. This is what came back.

0.39
Hazard ratio — four-gene mTOR composite in HCC

84.4 versus 37.8 months, from the published record alone.

The partitioned run returned two poles the monolithic query could not: renal cell carcinoma, where rapalogs succeeded, and hepatocellular carcinoma, where two decades of unselected trials did not. No patient data entered this step — the contrast came from the structure of the literature alone.

Asked whether that had a molecular correlate, a four-gene composite — high PPARGC1A with low EIF4E, FKBP1A and RRAGC — separated overall survival in public HCC cohorts by a median of 84.4 versus 37.8 months (HR 0.39, 95% CI 0.27–0.56; p = 1.1 × 10⁻⁷). No wet-lab experiments anywhere in it, and every candidate traces back to a source abstract.

Prognostic, not predictive. The signature has not been tested in rapalog-treated patients, and the composite was selected on the same cohorts in which the hazard ratio is reported. What it yields is a candidate patient-selection hypothesis for prospective evaluation — and an explanation for why the unselected trials failed. Manuscript submitted for publication.

The field's route

Two decades

EVOLVE-1546 patients on everolimus, hazard ratio 1.05. The trial did not test a drug so much as an average — everyone with the diagnosis, enrolled on the diagnosis.
Phase II through phase IIIRapalogs were taken deep into hepatocellular carcinoma without a selection marker, and the results were negative.
4–7 years, roughly $300MWhat a discovery half typically costs the industry, per programme — and the rapalog effort in liver cancer ran phase II through phase III on top of it.
The engine's route

One morning

4,453,833 abstracts, one runPartitioned into 23 disease corpora, one identical question to each, aggregation held to the end. The run itself: a morning of compute over the pre-built index.
Two poles, unpromptedRenal cell carcinoma, where rapalogs worked. Hepatocellular carcinoma, where they never did. No patient data entered this step.
Then the checkA four-gene composite, run afterwards against public HCC cohorts: hazard ratio 0.39, 84.4 versus 37.8 months, inside a subgroup the trials never defined. No wet lab, no new patients.

Nineteen immune contexts rejected the signature. One kept it. A signature that survives every context is usually measuring nothing — this one fails everywhere except an immune niche that is infiltrated but functionally restrained, which is where a real biological state should live.

Same disease, same drug class, same published record. The only thing that changed was how the corpus was asked. Nobody was withholding the answer — it was distributed across 4.45 million abstracts, and no one had read all of them at once. The index took months to build. The question took a morning to ask.

The Method

Partitioned Literature Aggregation.

Most AI answers a hard question by asking it once. Partitioned Literature Aggregation asks it hundreds of times instead. A question like what drives resistance to this drug is split into many smaller, biologically meaningful ones — which pathways predict survival, which mechanisms recur across independent datasets — each answered on its own evidence, then recombined into a ranked, weighted conclusion. The partitions come from folding the vector space so related concepts cluster and unlike ones pull apart; evidence for each is drawn from a bounded shell of that partition. Because every conclusion decomposes back into the evidence that produced it, a weak but reproducible signal survives instead of being averaged away by whatever is most cited.

Ingest → Partition → Evidence → Aggregate
01

Ingest

Index ~28M PubMed abstracts — the published biomedical record — plus internal corpora, refreshed on a rolling basis.

02

Partition

Fold the hyperdimensional space so related concepts cluster tightly and unrelated ones pull apart — producing partitions that are biologically meaningful, not merely topically near.

03

Evidence

Sample a bounded shell of each partition independently, clearing the phantom matches that look related but aren't. One question, one body of evidence.

04

Aggregate

Recombine the independent answers into a ranked, evidence-weighted result — the upstream gene, mechanism, or theme everything else moves around.

Search finds the needle. We find the magnet — the thing every needle points to.

From the Vector Space Day 2026 Keynote

The same method, on stage with Qdrant.

Slides from Saran Saund's VSD 2026 keynote in San Francisco, June 11, 2026 — the method above, walked through in front of a technical audience with the numbers attached.

Watch — VSD 2026 Shorts

Three minutes on the folding method.

Highlights from the VSD 2026 keynote posted to the Oncotelic YouTube channel. Vertical shorts optimized for phones — each landing on one core idea in under 60 seconds.

"1 in 10,000"

The attrition problem: 10,000 compounds enter development, one reaches approval. Why target selection is where it's decided.

"Scientists don't Google"

Why keyword search fails for hypothesis generation across the corpus.

"How we indexed 28M PubMed abstracts"

The pipeline behind the folding method — from raw abstracts to a searchable folded vector space on Qdrant.

Watch the full 20-minute VSD 2026 keynote — "Building the DNA of Search" with Bastian Hofmann (Qdrant), Saran Saund and Scott Myers (Oncotelic) — watch it on YouTube →

Partnership

Built on Qdrant. Widening the relationship.

Qdrant is the vector database layer underneath PDAOAI's knowledge engine. PDAOAI held a keynote slot at Qdrant's Vector Space Day 2026 in San Francisco on June 11, 2026, co-presented with Bastian Hofmann, Qdrant's Head of Product. Co-marketing and joint go-to-market are under discussion.

What folding buys any Qdrant customer.

Folding sits on top of Qdrant as a preprocessing layer. It doesn't replace vector search — it makes it precision-aware. The metrics a retrieval team tunes for — precision, search-space size, redundancy, downstream hallucination — are exactly the ones folding moves.

  • Smaller search space per query
  • Higher precision, lower redundancy
  • Scales to billions of vectors
  • Less downstream LLM hallucination
IP Position

Filed. Priority-dated. Patent-pending.

The stack is patent-pending. The lead application — covering the clustering method — is filed under the PCT with a favorable International Search Report. Folding, hybrid retrieval, and shell sampling sit in subsequent applications, filed or in preparation, and have not been examined. Nothing is granted.

WIPO's independent read on the core method.

Partitioned Literature Aggregation is the method filed as PCT/US2024/057972, which has drawn a favorable International Search Report and Written Opinion — all claims assessed as novel, inventive, and industrially applicable. The filing describes the method in claim language; PLA is the name we use for it. Manifold folding and probabilistic shell sampling sit in subsequent applications in the same priority family, filed or in preparation, and have not been examined. A separate application covering a discovery-side reasoning agent was filed prior to June 2026.

Peer-Reviewed Record

The engine, in the literature.

The method is documented in peer-reviewed publications, with the working stack disclosed in Methods — embedding, clustering and vector-database components named by version. In these papers the platform appears as “Oncotelic Chatbot technologies”, the name it carried at the time of writing.

Cancers · 2025

Positive Prognostic Overall Survival Impacts of Methylated TGFB2 and MGMT in Adult Glioblastoma Patients

Qazi S, Potts M, Myers S, Richardson S, Trieu V.
Cancers. 2025;17(7):1122

Engine in Methods §2.1 — AI-augmented literature analysis

doi:10.3390/cancers17071122 →

Int J Mol Sci · 2025

Comparative Tumor Microenvironment Analysis for HCC and PDAC Using KMplotter

Chang W-H, Shah D, Myers S, Potts M, Qazi S, Trieu V.
Int J Mol Sci. 2025;26(24):11920

Engine in Methods §4.1 — 32,264 abstracts processed

doi:10.3390/ijms262411920 →

Applied Work · Int J Mol Sci · 2025

Bioinformatic Approach to Identify Potential TGFB2-Dependent and Independent Prognostic Biomarkers for Ovarian Cancers Treated with Taxol

Qazi S, Richardson S, Potts M, Myers S, Saund S, De T, Trieu V.
Int J Mol Sci. 2025;26(24):11900

Conventional bioinformatics — this paper does not describe the engine. Included for the biology, which sits in the same taxane-treated setting as the case study above.

doi:10.3390/ijms262411900 →

All open access.

Deep Dive · In development

Quantum computing, on the part of the problem that is actually combinatorial.

Drug-like chemical space runs to roughly 10³⁰ molecules. Nobody enumerates that — you navigate it. The navigation is a combinatorial optimization problem, and that is the class of problem quantum hardware is genuinely suited to today.

What runs on quantum

Conformational search and candidate selection

Which substitutions, which conformers, which candidate set best satisfies the constraints against a fixed target. This is combinatorial optimization, not electronic structure — and we scope the claim to the former, not the latter.

How it plugs in

Target-first, not compound-first.

The upstream layers fix the target before any chemistry runs. That is what makes the optimization tractable — a bounded problem against one evidence-ranked target, rather than an open search across the whole space. Applied first to scaffolds that already carry human safety.

What this is not.

This is not quantum chemistry. The quantum layer solves optimization problems; it does not compute electronic structure, and no device does that at drug-relevant scale today. The claim here is narrow on purpose: the combinatorial layer of molecular design, against targets our own engine has surfaced and checked against patient-outcome data. In development.

Where the platform plugs in

Three things the method does today.

The stack applies to any corpus with a convergence problem. Three uses are running on internal programs now; evidence synthesis and clinical-insight workflows are the natural next ones.

01

Hypothesis generation

Point the engine at an open biological question and rank the convergence genes by literature sprawl. YAP1 → taxane resistance is the canonical example.

02

Repurposing & reformulation

Approved compounds mapped to newly surfaced convergence targets. The second compression axis: the engine cuts the timeline, an established safety profile cuts the uncertainty. Both PDAOAI programs work this way.

03

Target / pathway discovery

Cross-domain clustering identifies pathway hubs no single-paper reading would surface. Same method for oncology, neurology, fibrosis, neurodegeneration.

Related