The Engine

PDAOAI. Distill the record. Hunt the ghosts. Find the root drivers.

Distills 28M PubMed abstracts into one searchable substrate. Folds the vector space to make hidden structure visible. Ghost-hunts for convergence signals no single paper reveals. Quantum-driven optimization, in development, is being built to design against the root driver across the ~10³⁰ estimated size of drug-like chemical space.

Built on Qdrant
The Corpus
28M

PubMed abstracts — the totality of human biomedical knowledge.

The ~28M abstract-carrying citations indexed by the U.S. National Library of Medicine. Ingested, embedded in a Qdrant vector layer, folded to make convergence visible, and refreshed on a rolling basis as new science is indexed. This is the substrate PDAOAI builds on.

Case Study — VSD 2026 Keynote

How the engine recovered the convergence gene behind taxane resistance.

The worked example, step by step, with the slides from the keynote. Run as a positive control — YAP1's role is established in the literature, so the question had a knowable answer. What the corpus had to supply was the full picture across every domain it spans, which no single paper states.

YAP1
The convergence gene

2.5–4× more literature sprawl than any other candidate.

YAP1's literature footprint touches drug efflux, EMT, mTOR signaling, metabolism, and mechanotransduction — connections no single paper claims, but the structure of the corpus reveals. Reached from corpus structure alone, then checked against public patient-outcome data in taxane-treated ovarian cancer, held out of the ranking. Four genes, univariate, HR 1.17–1.45, unadjusted — an association within a treated population, not a treatment-interaction test.

The same method is being pointed at PDAC immune evasion, IPF fibrosis, and ALS neurodegeneration. One method, many workloads.

The Method

Folding the vector space.

Nobody can read 28 million abstracts, and the answer isn't in any one of them anyway — it's in the shape of all of them together. Ordinary vector search returns whatever sits closest to your query, which surfaces the most-studied gene rather than the one everything else depends on. PDAOAI's approach — what we call Structural Intelligence — folds the vector space with a learned transformation that clusters related concepts and pushes unlike ones apart, then samples a bounded shell of the manifold to clear the ghosts: the phantom matches that look related but aren't. Convergence becomes a measurable property.

Ingest → Fold → Sample → Discover
01

Ingest

Index ~28M PubMed abstracts — the published biomedical record — plus internal corpora, refreshed on a rolling basis.

02

Fold

Fold the hyperdimensional space so related concepts cluster tightly and unrelated ones pull apart.

03

Sample

Sample a bounded shell of the manifold, clearing the ghosts — the phantom matches that look related but aren't.

04

Discover

Surface the root driver: the upstream gene, mechanism, or theme everything else moves around.

Folding is the preprocessing step that makes Qdrant's search precision-aware — not just topic-aware.

From the Vector Space Day 2026 Keynote

The same method, on stage with Qdrant.

Slides from Saran Saund's VSD 2026 keynote in San Francisco, June 11, 2026 — the method above, walked through in front of a technical audience with the numbers attached.

Watch — VSD 2026 Shorts

Three minutes on the folding method.

Highlights from the VSD 2026 keynote posted to the Oncotelic YouTube channel. Vertical shorts optimized for phones — each landing on one core idea in under 60 seconds.

"1 in 10,000"

The attrition problem: 10,000 compounds enter development, one reaches approval. Why target selection is where it's decided.

"Scientists don't Google"

Why keyword search fails for hypothesis generation across the corpus.

"How we indexed 28M PubMed abstracts"

The pipeline behind the folding method — from raw abstracts to a searchable folded vector space on Qdrant.

Watch the full 20-minute VSD 2026 keynote — "Building the DNA of Search" with Bastian Hofmann (Qdrant), Saran Saund and Scott Myers (Oncotelic) — watch it on YouTube →

Partnership

Built on Qdrant. Widening the relationship.

Qdrant is the vector database layer underneath PDAOAI's knowledge engine. PDAOAI held a keynote slot at Qdrant's Vector Space Day 2026 in San Francisco on June 11, 2026, co-presented with Bastian Hofmann, Qdrant's Head of Product. Co-marketing and joint go-to-market are under discussion.

What folding buys any Qdrant customer.

Folding sits on top of Qdrant as a preprocessing layer. It doesn't replace vector search — it makes it precision-aware. Every metric a Qdrant customer cares about gets better.

  • Smaller search space per query
  • Higher precision, lower redundancy
  • Scales to billions of vectors
  • Less downstream LLM hallucination
IP Position

Filed. Priority-dated. Patent-pending.

The stack is patent-pending. The lead application — covering the clustering method — is filed under the PCT with a favorable International Search Report. Folding, hybrid retrieval, and shell sampling sit in subsequent applications, filed or in preparation, and have not been examined. Nothing is granted.

WIPO's independent read on the core method.

The lead application, PCT/US2024/057972, has drawn a favorable International Search Report and Written Opinion — all claims assessed as novel, inventive, and industrially applicable. Manifold folding and probabilistic shell sampling sit in subsequent applications in the same priority family, filed or in preparation, and have not been examined. A separate application covering a discovery-side reasoning agent was filed prior to June 2026.

Deep Dive · In development

Quantum computing, on the part of the problem that is actually combinatorial.

Drug-like chemical space runs to roughly 10³⁰ molecules. Nobody enumerates that — you navigate it. The navigation is a combinatorial optimization problem, and that is the class of problem quantum hardware is genuinely suited to today.

What runs on quantum

Conformational search and candidate selection

Which substitutions, which conformers, which candidate set best satisfies the constraints against a fixed target. This is combinatorial optimization, not electronic structure — and we scope the claim to the former, not the latter.

How it plugs in

Target-first, not compound-first.

The upstream layers fix the target before any chemistry runs. That is what makes the optimization tractable — a bounded problem against one validated target, rather than an open search across the whole space. Applied first to scaffolds that already carry human safety.

What this is not.

This is not quantum chemistry. The quantum layer solves optimization problems; it does not compute electronic structure, and no device does that at drug-relevant scale today. The claim here is narrow on purpose: the combinatorial layer of molecular design, against targets our own engine has already validated. In development.

Where the platform plugs in

Three things the method does today.

The stack applies to any corpus with a convergence problem. Three uses are running on internal programs now; evidence synthesis and clinical-insight workflows are the natural next ones.

01

Hypothesis generation

Point the engine at an open biological question and rank the convergence genes by literature sprawl. YAP1 → taxane resistance is the canonical example.

02

Repurposing & reformulation

Approved compounds mapped to newly surfaced convergence targets. The second compression axis: the engine cuts the timeline, an established safety profile cuts the uncertainty. Both PDAOAI programs work this way.

03

Target / pathway discovery

Cross-domain clustering identifies pathway hubs no single-paper reading would surface. Same method for oncology, neurology, fibrosis, neurodegeneration.

Related