Distills 28M PubMed abstracts into one searchable substrate. Folds the vector space to make hidden structure visible. Ghost-hunts for convergence signals no single paper reveals. Quantum-driven optimization, in development, is being built to design against the root driver across the ~10³⁰ estimated size of drug-like chemical space.
The ~28M abstract-carrying citations indexed by the U.S. National Library of Medicine. Ingested, embedded in a Qdrant vector layer, folded to make convergence visible, and refreshed on a rolling basis as new science is indexed. This is the substrate PDAOAI builds on.
The worked example, step by step, with the slides from the keynote. Run as a positive control — YAP1's role is established in the literature, so the question had a knowable answer. What the corpus had to supply was the full picture across every domain it spans, which no single paper states.
YAP1's literature footprint touches drug efflux, EMT, mTOR signaling, metabolism, and mechanotransduction — connections no single paper claims, but the structure of the corpus reveals. Reached from corpus structure alone, then checked against public patient-outcome data in taxane-treated ovarian cancer, held out of the ranking. Four genes, univariate, HR 1.17–1.45, unadjusted — an association within a treated population, not a treatment-interaction test.
The same method is being pointed at PDAC immune evasion, IPF fibrosis, and ALS neurodegeneration. One method, many workloads.
Nobody can read 28 million abstracts, and the answer isn't in any one of them anyway — it's in the shape of all of them together. Ordinary vector search returns whatever sits closest to your query, which surfaces the most-studied gene rather than the one everything else depends on. PDAOAI's approach — what we call Structural Intelligence — folds the vector space with a learned transformation that clusters related concepts and pushes unlike ones apart, then samples a bounded shell of the manifold to clear the ghosts: the phantom matches that look related but aren't. Convergence becomes a measurable property.
Index ~28M PubMed abstracts — the published biomedical record — plus internal corpora, refreshed on a rolling basis.
Fold the hyperdimensional space so related concepts cluster tightly and unrelated ones pull apart.
Sample a bounded shell of the manifold, clearing the ghosts — the phantom matches that look related but aren't.
Surface the root driver: the upstream gene, mechanism, or theme everything else moves around.
Folding is the preprocessing step that makes Qdrant's search precision-aware — not just topic-aware.
Slides from Saran Saund's VSD 2026 keynote in San Francisco, June 11, 2026 — the method above, walked through in front of a technical audience with the numbers attached.
Highlights from the VSD 2026 keynote posted to the Oncotelic YouTube channel. Vertical shorts optimized for phones — each landing on one core idea in under 60 seconds.
The attrition problem: 10,000 compounds enter development, one reaches approval. Why target selection is where it's decided.
Why keyword search fails for hypothesis generation across the corpus.
The pipeline behind the folding method — from raw abstracts to a searchable folded vector space on Qdrant.
Watch the full 20-minute VSD 2026 keynote — "Building the DNA of Search" with Bastian Hofmann (Qdrant), Saran Saund and Scott Myers (Oncotelic) — watch it on YouTube →
Qdrant is the vector database layer underneath PDAOAI's knowledge engine. PDAOAI held a keynote slot at Qdrant's Vector Space Day 2026 in San Francisco on June 11, 2026, co-presented with Bastian Hofmann, Qdrant's Head of Product. Co-marketing and joint go-to-market are under discussion.
Folding sits on top of Qdrant as a preprocessing layer. It doesn't replace vector search — it makes it precision-aware. Every metric a Qdrant customer cares about gets better.
The stack is patent-pending. The lead application — covering the clustering method — is filed under the PCT with a favorable International Search Report. Folding, hybrid retrieval, and shell sampling sit in subsequent applications, filed or in preparation, and have not been examined. Nothing is granted.
The lead application, PCT/US2024/057972, has drawn a favorable International Search Report and Written Opinion — all claims assessed as novel, inventive, and industrially applicable. Manifold folding and probabilistic shell sampling sit in subsequent applications in the same priority family, filed or in preparation, and have not been examined. A separate application covering a discovery-side reasoning agent was filed prior to June 2026.
Drug-like chemical space runs to roughly 10³⁰ molecules. Nobody enumerates that — you navigate it. The navigation is a combinatorial optimization problem, and that is the class of problem quantum hardware is genuinely suited to today.
Which substitutions, which conformers, which candidate set best satisfies the constraints against a fixed target. This is combinatorial optimization, not electronic structure — and we scope the claim to the former, not the latter.
The upstream layers fix the target before any chemistry runs. That is what makes the optimization tractable — a bounded problem against one validated target, rather than an open search across the whole space. Applied first to scaffolds that already carry human safety.
This is not quantum chemistry. The quantum layer solves optimization problems; it does not compute electronic structure, and no device does that at drug-relevant scale today. The claim here is narrow on purpose: the combinatorial layer of molecular design, against targets our own engine has already validated. In development.
The stack applies to any corpus with a convergence problem. Three uses are running on internal programs now; evidence synthesis and clinical-insight workflows are the natural next ones.
Point the engine at an open biological question and rank the convergence genes by literature sprawl. YAP1 → taxane resistance is the canonical example.
Approved compounds mapped to newly surfaced convergence targets. The second compression axis: the engine cuts the timeline, an established safety profile cuts the uncertainty. Both PDAOAI programs work this way.
Cross-domain clustering identifies pathway hubs no single-paper reading would surface. Same method for oncology, neurology, fibrosis, neurodegeneration.