Azati's hybrid retrieval architecture acts as a governed knowledge layer behind AI agents and search applications. Across 50M+ patent and biological-sequence records, a related patent-and-sequence intelligence engagement reported 70% less manual annotation and 90%+ search accuracy. On a narrower BLAST-tuning deployment for short DNA/RNA sequences, the same hybrid approach delivered 3x more relevant matches, cut false negatives by 55%, and accelerated analysis workflows by 40%.
Key takeaways
- Vector embeddings are built to capture meaning, not to reproduce exact strings. This makes them structurally weak on IDs, codes, and accession numbers.
- Patent and bioinformatics search is full of exactly those tokens: sequence identifiers, patent numbers, gene symbols, CAS numbers.
- In a related patent-and-sequence intelligence case, hybrid retrieval across 50M+ records reported 90%+ search accuracy and 70% less manual annotation.
- Hybrid retrieval, lexical (BM25) plus vector search, with re-ranking, is what closes that gap in production.
- In a related BLAST-based sequence-search engagement, hybrid tuning delivered 3x more relevant matches and cut false negatives by 55% for short DNA/RNA sequences.
- BLAST-based sequence alignment, LLM-assisted metadata enrichment, and a pluggable-LLM architecture round out the stack.
Why vector search alone underperforms on real queries
A vector embedding turns a document or a query into a point in high-dimensional space, and retrieval becomes a search for nearby points. That's genuinely useful for conceptual queries, "documents about reducing false positives in antibody screening", where no single keyword captures the intent.
This is a poor fit for a different, very common kind of query: one built around an exact token. As InfoQ notes in its analysis of hybrid retrieval for RAG, embeddings struggle with exact identifiers because an embedding model is trained to generalize meaning, and an ID is a string with no meaning to generalize.
Patent and biological-sequence search is dense with exactly this kind of token. A scientist searching prior art doesn't type "molecules similar to compound X". Instead, they type the CAS number. A bioinformatics analyst doesn't describe a gene in prose. They paste the accession number or the exact sequence fragment in ST.25 format. A patent examiner cross-referencing claims searches by publication number.
In each case, a vector index will return conceptually close documents and miss the one document that shares the exact string, because "close in embedding space" and "identical token" are different properties.
And we learned this the direct way. Azati’s first version of the retrieval layer for patent and sequence search leaned entirely on vector similarity, on the reasonable assumption that semantic search would generalize better than old-fashioned exact match. It looked strong in a demo built on conceptual queries. Yet, it failed the first time a real user typed an exact GenBank accession number and got back five plausible, wrong sequences instead of the one they'd asked for.
That failure got us thinking about moving the architecture toward an enterprise knowledge layer.
This is not an isolated failure mode. When AI agents tackle real-world tasks like reviewing patents, searching scanned engineering manuals, or checking compliance SOPs, relying only on vector search breaks down the moment specific details like part numbers, CAS codes, ticket IDs, or legal citations pop up.
An autonomous AI agent that retrieves "conceptually similar" documents instead of the exact identifier ends up with confident hallucinations. And closing this gap requires an enterprise knowledge layer. The one combining BM25 lexical indexing, vector search, and reranking to ensure agents receive strictly grounded context with verifiable source citations.
What hybrid retrieval actually does
Hybrid retrieval runs two mechanisms side by side to merge the results: a lexical engine (commonly BM25, via a search index like Elasticsearch) for exact and near-exact token matches, and a vector engine for conceptual similarity. A re-ranking step then scores the merged candidate set by relevance to the original query, so a strong lexical match on an ID and a strong semantic match on a described concept can both surface, instead of one silently losing to the other.
This is the same architecture Azati runs as a standing enterprise knowledge layer (see how it's structured for AI-agent workloads).
| Query type | Vector-only search | Lexical-only search (BM25) | Hybrid retrieval |
|---|---|---|---|
| Exact ID/accession number | Frequently misses the exact match | Reliable | Reliable |
| Conceptual/descriptive query | Reliable | Frequently misses relevant results using different wording | Reliable |
| Mixed query (ID + description) | Partial | Partial | Reliable |
| Typo or partial identifier | Unreliable | Unreliable without fuzzy matching | Reliable with tuned fuzzy matching |
IBM's overview of vector databases for RAG makes the complementary point from the infrastructure side: at billion-vector scale, approximate nearest-neighbor methods like HNSW are what make vector search fast enough to run in production at all. Which matters, because hybrid retrieval only works if both halves of the pipeline can run at the corpus's actual scale, not a demo-sized subset of it.
Neither half is optional at that scale: a lexical index without ANN-backed vector search stays fast but blind to meaning, and a vector index without a tuned lexical layer stays fast but blind to exact tokens.
Hybrid retrieval in patent and sequence search, in production and in proof-of-concept
For patent and biological-sequence intelligence, Azati's search core runs a hybrid pipeline. It combines BM25 for lexical and exact-match retrieval, vector search for conceptual similarity, and BLAST specifically for sequence alignment, because matching biological sequences is a unique challenge that standard search methods simply can't handle on their own.
We’ve proven that full-text hybrid search works across every part of a patent, well beyond just the claims, using the WIPO patent dataset as a working proof-of-concept. The same core pipeline using BM25 retrieval, AI query enrichment, and AI-generated summaries runs live in production today for sequence search.
On a related BLAST-based sequence-search engagement, we’ve tuned the algorithm's sensitivity thresholds and filtering logic for short DNA/RNA sequences under 20 base pairs. The result was three times more relevant matches, 55% fewer false negatives, and a 40% faster end-to-end analysis workflow.
Is hybrid agentic retrieval only relevant to patents and bioinformatics?
No. Patent and bioinformatics corpora represent the extreme stress-test of enterprise knowledge. The same architectural failure hits any enterprise environment where AI agents or employees operate across mixed data types:
- Engineering and intranets: Scanned equipment manuals, screenshots, or SOPs where part numbers and regulatory codes must be matched with zero error margin.
- Customer and partner portals: Technical support bases where users drop off after two failed searches and open a ticket instead.
- Regulated workflows: Legal, financial, and operational systems requiring air-gapped hosting (vLLM / self-hosted models), role-based access controls (RBAC), and explicit confidence thresholds.
A checklist for hybrid-readiness
Before shipping a RAG or enterprise search system, it's worth checking:
- Does the corpus contain IDs, codes, or identifiers users search by directly? If yes, vector-only retrieval will underperform on those queries specifically, not generally.
- Is there a lexical index (BM25 or equivalent) running alongside the vector index, or is vector similarity the only retrieval path?
- Is there a re-ranking step, or does the system just concatenate results from each retrieval method?
- Has retrieval accuracy been measured against the corpus's actual query mix, including ID-heavy queries, rather than only against conceptual test queries?
- Can the LLM component be swapped without redesigning the retrieval layer, for teams that need on-premises or self-hosted deployment later?
- Does the pipeline support grounded generation with source citations, allowing users and agents to trace answers directly to the original evidence?
- Is there an explicit confidence threshold and honest fallback mechanism so the system refuses off-topic queries or acknowledges missing data instead of fabricating answers?
- Can the retrieval core be deployed fully on-premises or within a self-hosted cloud tenant for zero-data-egress compliance?
What this means going forward
Semantic search at scale isn't a single-model problem. It's a retrieval-architecture problem, and the architecture has to account for the fact that real queries aren't uniformly conceptual. A meaningful share of them are exact, and no amount of embedding quality fixes that.
Hybrid retrieval isn't a workaround for an immature vector search. It's what production-grade retrieval looks like once a system has actually been tested against how people search, not just how a demo was scripted.
What’s really worth watching isn’t just bigger embedding models. It’s smart retrieval pipelines that combine lexical search, vectors, and domain tools like BLAST to make better decisions, rather than locking in a single approach and calling it a day.
Because, in a nutshell, production-ready retrieval is beyond a single algorithm. It's the underlying knowledge system that keeps AI accurate and grounded.