← /posts

Engineering

One search box was never enough: hybrid retrieval for company memory

Many teams building a company brain in 2026 still treat retrieval as a single vector-search problem. Here is why that fails exactly when reliability matters most, and how to build something better.

Author
Víctor Mollá
Published
Aug 7, 2026
Reading time
8 min
First on
X.com ↗

The problem nobody checks

You built semantic search over your company's documents. It answers questions. It is also wrong sometimes, with total confidence.

You assume the model is the whole problem. Often, it is not.

The problem may be the single lens you searched through. Pure vector search turns every question into a meaning-similarity match, which makes it vulnerable to queries that depend on exact terms, explicit relationships, effective dates, or source authority.

"What's the refund policy for enterprise clients?" is not only a semantic question. It also contains exact terms, "enterprise" and "refund", and an implicit requirement to retrieve the policy currently in force. A vector store can return the 2023 policy instead of the 2024 one because the two versions are semantically very close. Lexical retrieval, version metadata, active-document filtering, and reranking all have a role in preventing that error.

That silent failure, multiplied across every query that hinges on an exact identifier, a current version, or a governed relationship, is where trust in the system quietly leaks out.

One retrieval engine vs. several governed signals

A single retrieval pass is one lens on the truth:

SINGLE RETRIEVAL PASSquestionembeddingnearestneighbourstop-kanswerone method · one signal · one ranking

That is the atom: one method, one signal, one ranking.

Single retrieval methods optimize the signal they measure. A vector index is strong at semantic similarity but may underweight rare names or exact identifiers. A lexical index preserves term-level evidence but does not inherently understand paraphrases. A graph preserves explicit relationships but only for connections that have actually been modeled. Each signal has a different blind spot.

The graph's role is often misunderstood. A knowledge graph does not replace lexical or vector retrieval, and it is not inherently faster than either one. Its job is to preserve explicit structure that similarity search alone does not represent: which page mentions an entity, which claim concerns it, which sections belong to a page, and which pages explicitly link to one another.

Embeddings can make that graph semantically searchable. They help locate relevant sections, facts, or entities that become entry points into the graph. They do not decide that a factual relationship exists, and semantic proximity alone does not create a durable edge. Once candidate nodes have been found, graph traversal expands the evidence through explicit, named relationships.

This distinction is fundamental:

Embeddings answer "what appears relevant?" Graph edges answer "what is explicitly connected, and how?"

A traditional wiki already contains some of this structure through links, but those links depend on disciplined curation and can become inconsistent as the company grows. A governed property graph turns selected structure into queryable nodes and relationships, while the curated memory remains readable, versionable, and auditable.

A hybrid pipeline therefore combines complementary signals rather than asking one engine to correct every other engine in a fixed sequence. Lexical and vector retrieval find candidate evidence. Canonical lookup and claim retrieval add high-precision signals. The graph expands candidates through explicit relationships. Fusion and reranking decide which evidence is most relevant to the question.

For company-memory systems, this means one concrete shift: stop treating retrieval as "dump everything into a vector store and hope." Design the shape of retrieval first: what needs exact matching, what needs version filtering, what needs structural expansion, what needs semantic similarity, and when each authorization filter must run. Vectorizing documents alone was never enough; governance, permissions, provenance, and maintainability are part of retrieval quality.

From chunks to claims: why indexing is not governing

A knowledge base indexes documents. A Company Brain governs knowledge: it prioritizes active and authoritative evidence, records who said what, manages conflicts between sources, and decides when the available evidence is insufficient.

That governance operates across several layers.

Curated versionable memory

Every source is captured first for provenance and then integrated into a versionable, curated memory. This layer is not a copy of every raw transcript. It is a governed projection that preserves the facts, decisions, commitments, relationships, dates, and citations needed to answer questions. If that projection is wrong, the indexes built from it will be wrong too.

Chunk: the minimum retrieval and permission unit

Documents are divided into chunks or category-bounded segments. Each chunk carries its own permission scope and belongs to a governed knowledge category. Authorization is evaluated before evidence enters the ranking and retrieval pipeline, not after an answer has already been synthesized.

Claim: the atomic unit that gets governed

A curation or extraction model turns source material into isolated facts, metrics, commitments, decisions, and risks. Each claim retains temporal context, source provenance, author or speaker authority, confidence, and lifecycle state. That is what makes it possible to distinguish an active claim from a superseded one and to compare competing evidence explicitly.

The mistake many graph implementations make is treating "these two chunks talk about similar things" as a relationship by default:

  • "The chunk about customer X" and "the chunk about X's renewal"
  • "The chunk about pricing policy" and "the chunk about the volume discount"

Similarity is useful for proposing candidates, but it is not enough to create a factual edge. Before promoting a connection into the graph, ask:

  • Does the source assert a named relationship between grounded entities?
  • Does the document contain an explicit link to another governed page?
  • Does a claim identify a subject, object, owner, participant, dependency, or other meaningful role?
  • Does the structure itself justify the edge, such as a page containing a section or a fact being about an entity?

If yes, create a durable, named graph edge. The relationship remains meaningful even when the wording changes. If the only evidence is semantic proximity, keep it in the retrieval layer. It is a candidate signal, not graph truth.

Content hashes and checksums support incremental processing. A real change produces a different checksum, allowing the system to supersede and reindex only the affected unit instead of rebuilding the whole active projection. Unchanged content can remain untouched.

Your current "vector search over plain Markdown" may already behave like a primitive company memory. What it lacks is the ability to distinguish a missed term, an unmodeled relationship, an obsolete version, a denied permission, and a real absence of knowledge. Governance makes those states visible.

How this gets built in production

The pieces you need include:

  • A relational store with vector support for sources, documents, chunks, embeddings, claims, entities, and active projection generations.
  • A separate activity and audit capability. Separation is an architectural boundary.
  • A semantically searchable property knowledge graph as nodes connected through a small, explicit relationship vocabulary. Selected sections and facts carry embeddings so vector search and graph traversal can work together. Embeddings improve candidate discovery; they do not define the graph's factual topology.
  • A coordination layer that keeps background curation bounded and protects durable mutations.

The memory access layer exposes a small set of capabilities instead of one monolithic interface: initialize context, open an auditable activity, ask governed questions, discover candidates, retrieve authorized evidence, and request human clarification when the available evidence is insufficient or materially conflicted.

The retrieval flow is not "BM25, then graph, then vector." It is a composition of signals:

SIGNALS, COMBINEDquestionlexical searchvector searchcanonical lookupclaim retrievalpermissionsauthorizedchunks onlycandidatesgraphexpansionfusion+ rerankevidencefour signals, one ranking

Where a Company Brain actually breaks

Company Brain Engineering fails in predictable places. Know them before you hit them.

Permission leakage

If authorization exists only at the document level, a document containing both broadly available and restricted sections either exposes too much or denies too much. The fix is section- or chunk-level authorization evaluated before candidate ranking and evidence hydration. A chunk without permission should not enter the scoring pipeline.

Silent conflict resolution

When two sources contradict one another and the system does not detect or preserve that conflict, it effectively chooses a side inside an answer that sounds just as confident as if there were no disagreement. Claims need explicit temporal validity, source authority, author or speaker authority, lifecycle state, and contradiction links. Material conflicts should remain visible and be escalated to the relevant knowledge owner rather than silently resolved by retrieval rank.

Confusing memory with an execution engine

A technical stakeholder may compare the system to an agent that directly executes actions against operational databases and expect the Brain to do the same. That is the wrong boundary. The model consumes Company Memory for context, governance, evidence, and skills. The memory layer may capture knowledge and coordinate human clarification, but it does not become the general business-execution engine. Holding that boundary keeps knowledge governance auditable.

Confusing semantic retrieval with graph truth

Embeddings are probabilistic retrieval signals. Graph edges are durable assertions about structure or relationships. If the system converts nearest-neighbor similarity directly into factual edges, it creates a graph that is dense, difficult to explain, and expensive to correct.

The accepted trade-off is selective graph construction. Brain keeps a deliberately small relationship vocabulary and promotes only sufficiently grounded connections into durable edges. This sacrifices some automatic graph coverage, but it preserves interpretability, reduces semantic noise, and keeps traversals explainable. New relationship types can be introduced when the domain requires them; they are not invented per query by the embedding model.

Scaling to a living company-wide memory

Once the pipeline works on a handful of documents, scaling it to the whole company requires more than increasing corpus size. It requires controlled ingestion, versioned projections, incremental updates, explicit failure states, and observable quality gates.

In production, new source material is captured durably, curated through bounded background work, and incorporated into the current retrieval projection. Versioning and workload isolation keep updates controlled, while observable failure states make interrupted processing recoverable.

The retrieval and governance shape is:

RETRIEVAL AND GOVERNANCEcaptureprovenancecuratechunks, claimsindex+ graphexplicit edgesauthorizeper chunkretrievererankanswernot enoughevidence?ask apersonuncertain answers go to the person who knows

For organizations with stricter compliance requirements, the same logical architecture can be deployed with a different infrastructure topology. The retrieval model does not depend on whether the stores run in shared cloud infrastructure or a dedicated environment, although operational support, identity, networking, backups, and upgrades still need separate validation.

The curated memory layer should not hide uncertainty. It preserves evidence, active understanding, superseded claims, and unresolved contradictions. When the Brain does not have a reliable answer, it should say so, identify the missing or conflicting evidence, reach the person who can resolve it, and store the governed result so the same question does not require another hunt.

This is the shift Company Brain Engineering represents: stop treating retrieval as a contest over which search engine is best. Treat it as a layer-design and governance problem: what requires exact matching, what requires semantic retrieval, what deserves a durable graph edge, which source has authority, which evidence is permitted, and what a model must never decide on its own.

What changes when you think in layers instead of one search box

A pure vector store over thousands of documents has silent failure modes for exact identifiers, active versions, denied evidence, explicit relationships, and contradictory sources.

A governed pipeline combines lexical and vector candidate retrieval, explicit graph expansion, evidence reranking, authority-ranked claims, and human escalation. Its advantage is not that every failure disappears. Its advantage is that different failure states become distinguishable and actionable.

The graph contributes relational recall and explainability. Lexical and vector indexes contribute candidate recall. Claims contribute atomic governance. Versioning and authority determine what is current. Human escalation handles what the system cannot resolve safely.

That is not a marginal precision improvement. It is the difference between a company memory that merely sounds trustworthy and one that can explain why an answer deserves to be trusted.

The single search engine was never enough. The solution is not a graph made from semantic similarity; it is a governed memory in which similarity finds candidates and explicit relationships preserve meaning.

This is a technical breakdown of hybrid retrieval and governance patterns for company memory, as of August 2026. Stack details are illustrative of one real implementation: adapt the specific pieces (stores, reranking models, ingestion cadence) to your own environment before deploying at scale.

guest@gurusup > /posts

guest@gurusup > /install · connect Gurusup Brain to your agents

Get started - free