Skip to article content
    All articles

    Knowledge Graph vs Vector Database vs Ontology: Which Layer Decides AI Output Quality?

    Published . Updated .

    Most teams choosing an artificial intelligence (AI) foundation are asked to pick between a knowledge graph and a vector database. That framing is incomplete. Those two are storage and retrieval. The layer that decides whether the answers are correct is the ontology, and it is the one most projects skip.

    This guide defines all three, shows which failure each one causes, and gives the check to run before committing.

    Three layers: ontology (meaning) on top, knowledge graph and vector database (storage) below. How the three layers stack.

    What is a vector database, and what does it actually store?

    A vector database stores embeddings. An embedding is a list of numbers that represents the meaning of a piece of text, an image, or a record. Text with similar meaning produces numerically similar embeddings.

    When a question arrives, the database converts it to an embedding and returns the stored items whose embeddings sit closest to it. This is similarity search, and it is what most retrieval-augmented generation (RAG) systems run on.

    What a vector database is good at:

    • Finding relevant passages across unstructured documents nobody labelled
    • Tolerating paraphrase, so "cancel my plan" matches a page titled "ending your subscription"
    • Scaling to millions of documents with fast approximate lookups

    What it cannot do, by design:

    • Count or aggregate. "How many enterprise accounts renewed last quarter" is not a similarity question. The database returns passages that sound like the question, not an answer.
    • Follow a chain of facts. "Which customers of the accounts owned by this manager filed a support ticket about billing" needs three hops. Similarity has no hops.
    • Distinguish true from plausible. An outdated document and a current one embed almost identically. The database has no notion of which is correct.
    • Enforce permissions natively. Access control has to be bolted on at the application layer, and that is where leaks happen.

    What is a knowledge graph, and how is it different?

    A knowledge graph stores entities as nodes and the relationships between them as typed edges. A customer node connects to an order node through a "placed" edge. The order connects to a product. The product connects to a supplier.

    Where a vector database returns text that resembles the question, a graph returns facts that are connected to it. The three-hop query above is a traversal, and traversal is what graphs are built for.

    What a knowledge graph is good at:

    • Multi-hop questions where the answer sits two or three relationships away
    • Exact joins, counts, and aggregations over structured business objects
    • Provenance, because an edge can carry where the fact came from and when
    • Permissions, because access can be enforced at the node and edge level

    Where it struggles:

    • Unstructured content. A graph cannot ingest a folder of PDF files and become useful. Something has to extract the entities first.
    • Coverage. A graph only knows what was modelled. Anything outside the schema is invisible rather than merely hard to find.
    • Cost of change. Adding a new entity type is a modelling decision, not an upload.

    What is an ontology, and why is it not the same as a knowledge graph?

    This is the distinction that most comparisons miss, and it is the one that decides output quality.

    A knowledge graph is the data. An ontology is the agreement about what that data means.

    The ontology defines which entity types exist in the business, which relationships between them are valid, and what each term denotes. It is the answer to questions like:

    • What counts as an "active customer"? Logged in within 30 days, or holding a paid contract?
    • Is a "renewal" a new contract or a continuation of the old one?
    • Can an employee belong to two departments at once?
    • When Sales says "account" and Finance says "account", are those the same object?

    A graph can be built without settling any of this. Most are. The result is a system that returns a confident number that two departments dispute, which is worse than returning nothing.

    The practical test: if two agents query the same graph and give two different answers to "how many active customers do we have", the problem is not the database. The ontology was never agreed.

    Chiri Field Guide No. 01. A Field Guide to Ontology covers what an ontology is in plain terms, how your team builds one from your own live data, why it decides whether AI scales or stalls, and what it protects once you own it. Read the field guide: /field-guides/ontology

    How do the three compare?

    Vector database Knowledge graph Ontology
    What it holds Embeddings of content Entities and relationships Definitions and rules
    Question it answers What text is similar to this? What is connected to this? What does this term mean here?
    Good at Recall over unstructured content Precision, multi-hop, aggregation Consistency across teams and agents
    Fails at Counting, reasoning, truth Unmodelled content Nothing on its own. It is not a store.
    Failure you see Plausible answers that are wrong "No results" on valid questions Two correct answers that disagree
    Who usually owns it Engineering Data platform Nobody, which is the problem

    Evaluating this against two other options? That is the part most teams get wrong, and it is the work Chiri does before anything gets built. Talk to our team: /contact, or see how Chiri Brain handles it: /platform.

    Which one do you actually need?

    The honest answer depends on the question your business asks most often.

    Start with a vector database if your highest-value use case is helping people find things inside documents nobody has structured. Support deflection, policy lookup, and internal search all fit. The content already exists and nobody is going to model it.

    Start with a knowledge graph if the questions that matter are about relationships between business objects that already live in systems of record. Revenue, entitlements, org structure, supply chain, and risk exposure all fit. The data is structured and the questions have exact answers.

    Settle the ontology first in either case if more than one team will ask the same question, or an agent will act on the answer rather than show it to a human who can sanity check it. An agent that books, refunds, or escalates on a disputed definition creates a liability, not a productivity gain.

    Most mid-market deployments end up needing all three. The sequencing matters more than the selection.

    What breaks when you pick the wrong one?

    Three failure patterns show up repeatedly in stalled pilots.

    The confident wrong number. A team ships a RAG assistant over finance documents. Asked for quarterly revenue, it retrieves a passage from a superseded forecast and reports it as fact. Similarity search cannot tell a draft from a final. The fix is not a better model.

    The agent that cannot answer a simple question. A graph is built from the customer relationship management (CRM) system only. A question spanning CRM and support tickets returns nothing, because support was never modelled. Users conclude the system does not work and stop asking.

    The two dashboards that disagree. Two agents, both technically correct, report different headcounts because one counts contractors. Trust collapses faster from this than from an outright error, because there is nothing to fix in the code.

    How do the three work together?

    The layers stack rather than compete.

    1. The ontology defines the vocabulary. Entity types, valid relationships, and the meaning of each term, agreed across the teams that will use it.
    2. The knowledge graph instantiates it. Real entities and relationships, populated from systems of record, carrying provenance and permissions.
    3. The vector index covers the rest. Unstructured content that was never modelled, retrieved by similarity and grounded against the graph before it reaches the user.

    The routing rule is straightforward. If the question has an exact answer, the graph answers it. If it needs judgement over prose, the vector index supplies passages and the graph supplies the entities those passages are about. The ontology keeps both honest about what the words mean.

    What should you check before you build?

    Five questions, answerable in a workshop rather than a proof of concept:

    1. Write down your ten most valuable questions. Count how many have an exact answer. That ratio tells you the graph-to-vector balance.
    2. Take three terms every team uses and ask four teams to define them. If the definitions differ, the ontology is the first piece of work, not the last.
    3. Ask where the answer would come from today. If a human would open two systems, the graph needs both. Modelling one is the common shortcut and the common failure.
    4. Decide who can revise a definition. An ontology without an owner drifts back into disagreement within two quarters.
    5. Decide whether an agent acts or advises. Advising tolerates ambiguity. Acting does not.

    Want the long version? The Field Guide to Ontology walks through building one from your own systems, with the failure modes above worked through in detail. Open it here: /field-guides/ontology

    Questions and answers

    Is an ontology just a database schema? No. A schema says a column is a string. An ontology says what the string means, which values are valid, and how the object relates to others in the business. A schema constrains storage. An ontology constrains meaning.

    Can you use a vector database without a knowledge graph? Yes, and it is the right starting point when the use case is search over unstructured documents with a human reading the result. It becomes the wrong choice as soon as answers need to be exact or an agent acts on them.

    Do you need a graph database to have a knowledge graph? No. Graphs can be implemented on relational stores. The choice of store matters less than whether the entities and relationships are modelled and maintained.

    Is a knowledge graph the same as an ontology? No. The ontology is the model. The knowledge graph is the populated instance of that model. One ontology can describe many graphs.

    Which comes first if the budget only covers one? The ontology, because it is the cheapest to produce and the most expensive to retrofit. It is a set of agreed definitions rather than infrastructure, and every later choice depends on it.

    How long does defining an ontology take? For a single high-value workflow at a company of 50 to 500 employees, the first usable version is a matter of weeks, not quarters. Attempting to model the whole business at once is the reason ontology projects get a reputation for never finishing.