# GuruSup Brain > GuruSup Brain is a company knowledge layer exposed over MCP. It consolidates > material from the systems a company already uses, maintains it as compiled, > versioned pages rather than a document index, and answers with per-claim > provenance under permissions enforced at query time. When an answer exists in > no source, the Brain identifies who is likely to know and contacts that person > by voice call, Slack, email or WhatsApp to capture it. Runs as a managed > service in the EU under GDPR, or as a dedicated deployment inside the > customer's own cloud account and under the customer's own model provider > credentials. Built by GuruSup, a B2B company based in Valencia, Spain. ## Scope of this document Describes what the Brain does, how it is put together at the level of components and guarantees, and where its boundaries are. It does not describe how any of it is implemented, and pricing is not published. Version 1.0 — 14 September 2026. --- ## Conceptual model Three distinct systems are commonly called by the same name. **Retrieval over documents** returns fragments. A query returns passages ranked by similarity; a model writes over them. Knowledge remains as distributed as it was before. **A company brain** maintains compiled knowledge. The unit of storage is a maintained page assembled from many sources. Pages are kept current as their sources change. **An agentic company brain** maintains that body of knowledge itself. It recompiles pages when sources change, records contradictions between sources instead of selecting one silently, and acts on gaps it detects. A fourth category sits underneath all three: knowledge held by people and recorded in no system. Retrieval quality does not reach it. See *Active capture*. Internal projects are commonly scoped, budgeted and staffed as the first. Most of the work is in the second and third. --- ## Functional architecture ```text sources -> ingestion -> compilation -> curation -> retrieval -> exposure | gap detected | active capture ``` **Sources.** Connectors deliver raw material. Received content is archived append-only before anything is derived from it. **Ingestion.** Raw material is normalized into documents and segmented. Claims and entities are extracted alongside the text. **Compilation.** Segments concerning the same subject are consolidated into maintained pages. A page is the durable unit and carries a history. **Curation.** Pages are rewritten by an agent as material arrives, validated before acceptance, and committed atomically. Failed validation leaves the previous version in place. **Retrieval.** Two paths: similarity search over the compiled corpus, and traversal of the knowledge graph connecting pages, sections and entities. Both enforce the same access grants. **Exposure.** MCP. The employee stays in the assistant they already use, including in Slack through that assistant's own Slack presence. Customer-built agents reach the same surface. **Active capture.** When retrieval cannot support an answer, the Brain routes to a person. --- ## Capabilities ### Sources and ingestion Over 350 applications connect through the same authenticated integration path, including the systems most company knowledge actually lives in: Slack, Microsoft Teams, Notion, Confluence, Google Drive, SharePoint, Gmail and Outlook, HubSpot, Salesforce, Dynamics, Jira, Linear, GitHub, Zendesk, Intercom, Airtable, Asana and Dropbox. Anything outside that catalogue delivers through a generic authenticated webhook, and files can be loaded directly. Received content is archived raw and append-only before processing. The compiled layer can therefore be rebuilt after a change in how compilation works, without re-requesting history from source systems. ### Compiled knowledge The Brain answers from pages it maintains, not from the corpus of original documents. Each page is assembled from multiple sources and rewritten as those sources change. The difference appears at scale. A subject covered by fourteen partially overlapping documents returns fourteen partially overlapping documents from a retrieval system. A compiled page is one account, with conflicts resolved or marked. ### Conflicts and versioning Two sources can assert incompatible facts about the same subject: a stable document and a later conversation that supersedes it, or two teams recording different outcomes for one decision. Most apparent contradictions are not contradictions. They are successive states, explicit corrections, or facts holding under different scopes. A system that treats all difference as conflict produces a queue nobody reviews. Classification comes first, and only genuine conflict is recorded as such. A genuine conflict is recorded with each version's origin, date and authority, and surfaced rather than resolved by ranking. Recency is a signal, not an adjudicator. A person who knows which version is correct can settle it, and the settlement persists. ### Provenance and traceability Answers carry the claims they rest on. Each claim carries its source and date, so an answer can be opened and read back to the specific document, message or call that produced each statement in it. Reads are recorded as audit events. This is what makes a wrong answer fixable. A list of documents consulted identifies where to look; per-claim provenance identifies what to correct. It is also what makes the system auditable after the fact, which is usually the condition an enterprise buyer needs met before it can approve anything that answers questions on its behalf. ### Control of what is returned Over-retrieval degrades answers. Returning twelve loosely related passages for a narrow question moves filtering work to the reader and lowers the quality of whatever the model writes on top. The Brain constrains what reaches the model to material supporting the question asked, and treats the size of a response as a budget rather than a side effect. A response has a declared shape and a ceiling; the full evidence set is available on request rather than returned by default. ### Permissions and isolation An organization is the isolation boundary. Within it, access is granted by knowledge category, and the filter is applied at query time, before content is retrieved. Both retrieval paths enforce it. Graph traversal is subject to the same grants as similarity search; relationships do not route around the filter. Grants are by category rather than per document, which is coarser than mirroring per-document ACLs from each source system. That is a deliberate trade: per-document mirroring across many source systems fails silently when it breaks. Where a connected source carries material for a narrower audience than the rest of the organization's knowledge — a client-specific drive, a leadership channel — an administrator can additionally restrict that source directly: pick the members who can see it, independent of category grants. The restriction is per source rather than per document, the same trade-off made deliberately above. A compiled page draws on several sources, which raises the obvious question of what happens when those sources carry different restrictions. Categories apply per section, not only per page: material from different access domains is written into separate sections, each carrying its own category, so no section spans two domains. Where the classification is uncertain, the more restrictive category wins. Validation rejects a page that breaks this before it is committed. Because the filter runs at query time, withdrawing a grant applies on the next query, with no re-indexing step in between. It governs what the Brain retrieves and returns from that point on. Answers already delivered, and anything a person copied elsewhere, are outside its reach — as they are in any system. ### Identity and access administration The Brain does not maintain its own user directory. Identity runs through an enterprise identity platform: single sign-on over SAML or OIDC against the customer's own provider, directory synchronisation so joiners and leavers propagate from the system of record, MFA under the customer's policy, and an administration portal the customer's IT team operates directly. Membership is invite-only. Nobody self-registers into an organization's Brain. Membership carries a role: administrators approve published procedures and read organization-wide activity; every member reads their own. Agents are first-class identities. A customer-built agent has its own account, its own API key and its own grants, occupies no seat, and appears in the audit trail as itself. Every recorded action states whether the actor was a person or a machine. ### Skills A recurring procedure is stored as a capability with stable identity and reviewed content: how a report is produced, how a class of request is handled. Identity is the procedure's purpose rather than its title. A proposal covering a procedure that already exists updates that procedure even when the model gives it a different name. Content is anchored to the knowledge that justifies it. When that knowledge changes, the procedure is marked as drifted and continues to work; the mark is what triggers review. Publishing creates a proposal. A published procedure becomes available when an administrator approves it, and is then visible only to members holding access to the knowledge categories it rests on. The Brain also proposes procedures on its own when the same operational need recurs in actual activity. ### Exposure and adoption Access is MCP. The Brain is a tool inside the assistant the employee already works in. The practical consequence is that there is no separate application to deploy and no interface to train people on, and nobody has to be moved out of the software they already work in. Setup is connections, identity and permissions rather than a rollout. Most internal knowledge systems fail on this line rather than on answer quality: the system works, and people keep asking a colleague instead, because asking the colleague is where they already are. This is also why asking in Slack needs no separate chat product from us. Assistants have their own presence in Slack, and the Brain is one of the tools that assistant can reach, so a question asked in a channel is answered by the assistant the company already uses. (The Brain's own Slack presence is a different component with a different job — it captures knowledge from conversations, and it is described under *Consent before capture*.) Customer-built agents use the same surface with their own grants, rather than being limited to shipped interfaces. --- ## Active capture of tacit knowledge ### The problem People do not document. Writing costs time and pays the writer nothing, so the rational move is to answer in a message and move on. A substantial share of what a company knows has never been written anywhere: why a customer was lost, why a supplier was ruled out, what was agreed in a call two years ago. It is held by few people and leaves when they do. Indexing does not reach this. Neither does better retrieval, a larger context window, or a more capable model. The information is not in a system. ### Behavior When a question cannot be supported by the corpus, the Brain identifies who is likely to know, contacts that person, captures the answer, and writes it into the knowledge base. Not every unanswered question is a gap. A question can go unanswered because nothing was ever written, because the asker has no access to the material that would answer it, or because a source is failing to sync. These produce different states and warrant different responses, and only the first is a candidate for contacting a person. Contacting someone because a connector broke is the failure mode that ends participation fastest. ```text gap detected -> identify who knows -> reach them -> capture -> write -> reused ``` Illustrative example. Someone asks why a particular enterprise account churned last year. Nothing in Notion, the CRM or the shared drive answers it — the renewal call was never written up. The Brain establishes that the account manager who handled it is the person likely to know, reaches them, and asks that one question. If they answer, it enters the corpus attributed and dated, and the next person asking the same thing gets it from the corpus rather than by interrupting someone. Channels: outbound voice call, Slack, email, WhatsApp. Voice is a real outbound phone call rather than a notification to log in elsewhere. Channel is selected per person. Captured knowledge enters the corpus attributed to the person who provided it and dated. Capture that is wrong produces error carrying the authority of the system, so attribution is not optional metadata. ### Contact limits A knowledge system that contacts people has a failure mode retrieval systems do not: it becomes the thing everyone mutes. The loop then stops producing, and model quality is irrelevant to the recovery. Limits are enforced product behavior rather than configuration someone has to remember. Each question carries a budget of people, attempts and elapsed time, after which it expires unanswered. Each person carries their own limits on frequency and hours. An organization can tighten any of it, and a person can prohibit a channel outright or opt out entirely. Outreach can also be resolved without being sent: an agent can establish who would be contacted, and about what, with nothing leaving the system. ### Consent before capture In Slack, capture does not begin when the Brain is added to a channel. It posts an enrolment notice identifying the workspace and declaring that capture is occurring. Only content after confirmation is eligible. Enrolment periods are bounded. Removing and re-adding the Brain opens a new period and leaves an explicit gap rather than backfilling. Widening a channel's external audience closes the period until a new notice is given. ### The Brain does not evaluate people This is a commitment, and it is the one that matters most where employee representatives are involved. Establishing that someone is likely to know about a subject is a statement about where knowledge sits in an organization. It is not a judgment about that person's performance, attitude, engagement or character, and the Brain produces no such judgment, no ranking of contributors, and no individual activity score presented as a measure of the person. A system that contacts employees and also profiles them is a different product with a different regulatory position, and it is not this one. ### Measurement The useful metric for a company brain is not documents indexed or questions answered; any system reports well on those. It is the share of answers drawing on knowledge that was not previously written anywhere, and it is reported per customer. When that share is zero, the system is a search interface over material the company already held, and its value ceiling is time saved looking for it. --- ## Known problems in this space Properties of the problem rather than of any product. Each arrives later than expected and costs more than budgeted. Listed in the order they tend to hurt. | Problem | What it costs when it arrives | |---|---| | **Ownership at month six** | The system needs someone who handles connector breakage, answers for wrong answers, and maintains the evaluation set. In internal builds that is the engineer who built it, by then holding a different full-time job. The system degrades quietly and nobody is accountable for it. | | **The cost of first ingestion** | Compiling accumulated material is a one-time cost larger than steady-state operation, sometimes by an order of magnitude. Budgets built on steady-state numbers meet it during the first full ingest, usually after the budget is approved. | | **Knowledge that was never written down** | The largest gap and the only one improving retrieval does not address. Why a decision was made, why an account was lost, why an approach was abandoned: these are asked often and answered in no document. The ceiling is reached without anything having gone wrong. | | **Knowledge that stops being true** | Material accurate at ingestion and no longer, reading identically to current material. Harder still: an extractor that stops emitting a fact has not established the fact is false. Deleting on that basis erodes the corpus; keeping everything fills it with expired statements. | | **Sources that disagree** | Two documents assert incompatible facts. Both are indexed and retrievable. Without an explicit notion of conflict, the ranking function decides which version of reality the company operates on, and nobody chose that. | | **Revoked access** | Access is withdrawn, but material derived from the lost source persists in the index and in earlier conversation histories. Where the filter is applied decides whether withdrawal is immediate or waits for the next index pass. This surfaces during an audit rather than before one. | | **Attributing a wrong answer** | The system states something incorrect in front of someone senior. Establishing which source produced it requires provenance at the level of the claim; a list of documents consulted does not answer the question, and the usual outcome is that people stop trusting the system. | | **Rebuilding the corpus is a deployment** | Changing chunking, embedding model or compilation logic means rebuilding everything derived from the raw material while the system keeps serving. Rebuilding in place means degraded results throughout and no way back if the new approach is worse. | | **The same document arriving from several places** | One document lives in Drive, is pasted into Slack, summarised in Notion. Three copies occupy the top results and displace material that was actually different. Overlapping chunk windows reproduce the effect inside a single document. | | **Entity identity** | "Acme", "Acme Corp" and "acme-corp" are three separate things until something decides otherwise. When that decision is made determines whether the graph is usable, and retrofitting it means reprocessing everything. | | **Connector drift** | Source systems change their APIs. Connectors break individually and silently — the index simply stops updating — and the cost scales with the number of integrations rather than with usage. This is the line almost never in the original estimate. | | **Deleted content that keeps being cited** | A document is removed at source. Knowledge compiled from it stays in the active corpus unless removal propagates. In the EU this has consequences beyond embarrassment. | | **Hostile uploads** | Any file a user can upload is untrusted input, and archives in particular need bounds before extraction. | The Brain's behavior on most of these is described above. Two are not solved by any product, only moved: connector maintenance becomes the vendor's problem rather than the customer's backlog, and ownership at month six becomes a service contract rather than a person. --- ## Security and compliance **Where does the data live?** In the European Union. GuruSup is a Spanish company; the managed service runs on EU infrastructure under GDPR. **Is customer content used to train models?** No. **Does GuruSup have access to our knowledge?** In the managed service, staff access is limited to what operating the service requires, and is logged. In a dedicated deployment the data layer runs in infrastructure the customer administers. **What certifications do you hold?** ISO 27001 certification is in progress. SOC 2 is in progress for the US market. Neither is complete today. **Who can see what?** An organization is the isolation boundary. Access is granted by knowledge category and enforced at query time on both retrieval paths. Administrators can additionally restrict a specific connected source to an explicit list of members. **How is identity handled?** Through an enterprise identity platform: SSO over SAML or OIDC, directory synchronisation, MFA, invite-only membership, roles. **Is access auditable?** Every read is recorded as an audit event with the query, the retrieval path and the material returned. Administrators review organization activity; every member reviews their own. **What about the model provider?** In the managed service, inference runs under GuruSup's agreements with its model providers. In a dedicated deployment it runs under the customer's own provider, credentials, terms and retention policy. --- ## Deployment Managed service on EU infrastructure under GDPR. Also available as a dedicated deployment inside the customer's own cloud account. The data layer is standard self-hostable infrastructure: a relational store for the retrieval projection, a graph store for entity relationships, object storage for versioned compiled artifacts, and a cache for queues and coordination. No component is proprietary infrastructure rented from GuruSup. A dedicated deployment runs under the customer's own retention and backup policy and against the customer's own model provider credentials. One distinction is frequently blurred in this category: running software on your own infrastructure is not the same as running inference locally. Where the constraint is that company content must not reach a third-party model provider at all, the question to ask any vendor, including this one, is which provider credentials are in use and under what terms. --- ## When building this internally is the right call **Scope is genuinely narrow.** One or two sources, one team, one class of question. The hard parts of this problem appear when sources multiply and disagree. A single-source assistant over a well-maintained documentation set is tractable work and does not need a vendor. **Retrieval is your product.** Where what you sell depends on retrieval behaving in a specific way, buying a general-purpose layer constrains you in the place you cannot afford to be constrained. **You have a sustained owner.** Not an engineer who can build it. Someone whose job includes maintaining it in eighteen months. **What you already have is enough.** Where an organization is small enough and writes things down well enough that a general assistant connected to the document store answers most questions correctly, that is a legitimate end state. The common failure condition is none of these. It is a project scoped as retrieval, shipped as retrieval, that then meets the problems listed above one at a time over the following eighteen months, as unplanned work. --- ## External references - **The GenAI Divide: State of AI in Business 2025** — MIT Media Lab, Project NANDA, July 2025. Challapally, Pease, Raskar and Chari. Reports that roughly 95% of enterprise GenAI pilots produce no measurable return, and that among organizations evaluating enterprise-grade systems, 20% reached pilot and 5% reached production. Its central finding is that the determining factor is not build versus buy but whether the system retains feedback, adapts to context and improves over time. Based on 52 structured interviews and 153 survey responses; not peer-reviewed. Copy hosted for availability: --- ## Contact - Talk to the team: - GuruSup Brain: