14 LAUNCHES IN 14 DAYS

SEE WHAT’S NEW

AI Agent Data Model: How to Give Every Answer One Definition

AI Agent Data Model: How to Give Every Answer One Definition

AI Agent Data Model: How to Give Every Answer One Definition

All Posts

What an AI agent data model needs: column descriptions, keys, live values and versioned metric definitions, with the research behind each part.

Ka Ling Wu

Co-Founder & CEO, Upsolve AI

10 min

Title card for Upsolve Data Models, Feature 01 of 14 in The Complete Agent Stack

An AI agent data model is the registered description of your tables, columns, keys, live values and metric definitions that an agent reads before it writes a single query.

This guide is part of The Complete Agent Stack series on the fourteen components a production data agent needs. Part 01 starts at the bottom of the stack.

Key Takeaways

  • Business context can matter more than model choice. On the BIRD text-to-SQL benchmark, GPT-4's test-set execution accuracy rose from 34.88% to 54.89% when each question came with a short piece of external knowledge about the data, and ChatGPT with that knowledge (39.30%) outscored GPT-4 without it (Li et al., NeurIPS 2023).

  • Column descriptions are measurable context. In a study built on BIRD, GPT-4o's execution accuracy went from 30.13% with no column descriptions to 36.78% with generated descriptions on the BIRD development set, and the authors report gains of over 20% for some models when column names were made completely uninformative (Wretblad et al.).

  • Live column values close a gap that schemas leave open. BIRD lists "value illustration" (the types and categories of values a column holds) as one of four kinds of knowledge models need, which is why an agent should know every contract status in your data before it filters on one.

  • Keys and relationships prevent silent double counting. dbt's MetricFlow uses declared entities as join keys and refuses fan-out joins, the joins that turn one row into several and inflate a sum (dbt docs).

The 45-second version:

What Is an AI Agent Data Model?

An AI agent data model is a structured, registered description of the data an agent is allowed to query. It lists the tables and fields the agent can see, explains each one in plain language, records primary and foreign keys, and stores the real values that categorical columns contain. Paired with a set of canonical metric definitions, it gives the agent the same orientation a senior analyst would give a new hire on day one.

Diagram of The Complete Agent Stack: four layers and 14 features, with Data Models highlighted in the ontology layer

The term overlaps with three older ideas. A data dictionary documents columns for people. A semantic layer defines metrics so that every BI tool computes them the same way, which we cover in depth in Semantic Layer for AI: Why Data Agents Hallucinate Numbers. An ontology describes the entities in a business and how they relate. An agent data model borrows from all three, but its reader is a language model that will write SQL, so every field in it exists to change what that SQL looks like.

A person reading a data dictionary can ask a colleague or notice that a number looks odd; an agent can do neither unless its context prompts it. Anthropic's engineering team describes the job as finding "the smallest set of high-signal tokens" that make the desired outcome likely (Anthropic). A data model is that set of tokens for structured data.

Why Do AI Agents Pick the Wrong Metric Definition?

Most companies carry several working definitions of the same metric. Revenue can be booked, billed or recognised, and ARR can include or exclude usage and trials. Each definition is correct for some team, and the schema rarely says which one applies to the question in front of the agent.

A database schema records structure. It says that a table called contracts has a column called status of type text. It does not say that finance counts only active and renewing contracts as revenue, or that invoices_v2 replaced invoices while a legacy job still fills the old table. Without that knowledge, the model chooses whichever interpretation seems most plausible from the column names.

The BIRD benchmark measured this gap directly. Its authors built 12,751 question and SQL pairs over 95 databases totalling 33.4 GB, and annotated questions with "external knowledge evidence", such as a domain rule or an explanation of a coded value (Li et al.). With that evidence, GPT-4's test accuracy rose by 20.01 points. Even then, the best model sat far below the 92.96% that human annotators reached.

The failure mode is what makes this expensive. An agent without a definition rarely refuses; it writes valid SQL against a reasonable interpretation and returns a specific number. The number looks right, and nobody checks it until it disagrees with finance's figure in a board pack.

What Belongs in a Data Model for AI Agents?

A useful data model has five parts. Each one answers a question the agent would otherwise guess at.

Table and Column Descriptions

Descriptions tell the agent what a table represents, which table is the source of truth for a concept, and what a column actually holds. A column called amt could be gross, net, pre-tax or in cents. One sentence removes the ambiguity.

Wretblad and colleagues generated and hand-refined column descriptions for the BIRD development set, then tested several models with and without them (Wretblad et al.). Descriptions consistently improved accuracy, particularly for larger models such as GPT-4o. The authors also found that models struggle to describe columns that are genuinely ambiguous, and recommend that human experts annotate those columns while models draft the obvious ones. In practice, that means generating the easy majority and spending expert time on the columns where teams disagree.

Primary Keys, Foreign Keys and Relationships

Keys tell the agent how tables join and what one row represents. Without them, an agent can join orders to line items and sum order totals across every line, producing a figure several times too large with no error raised.

Semantic-layer tools treat this as a hard constraint: dbt's MetricFlow builds a graph from declared entities and "restricts the use of fan-out and chasm joins" so that aggregations stay correct (dbt docs). Snowflake's semantic view specification similarly asks for a primary key per table and infers relationship types from the keys and the data (Snowflake docs). In prompt-based text-to-SQL, listing foreign keys raised accuracy by 0.6 to 2.9 points for most OpenAI model and prompt combinations in one benchmark study (Gao et al.), a small average that hides the questions where a wrong join multiplies the answer.

Selectable Column Values

Many other wrong answers come from filters. A user asks for "open contracts", the agent writes WHERE status = 'open', and the column actually holds Active, Pending Renewal, Expired, Terminated and Draft. The query returns zero rows, or the agent quietly picks the wrong subset.

The fix is to give the agent the real values of low-cardinality columns before it writes SQL. Snowflake's specification tells authors to add any sample value "likely to be referenced in the user questions" (Snowflake docs). Research systems go further: the CHESS framework retrieves relevant database values as a dedicated step, and its schema selector cut the tokens sent to the model by a factor of five while improving accuracy by about 2% (Talaei et al.).

Metric Definitions and Business Vocabulary

Some context belongs to the business as a whole and has no natural home in a column description. What counts as an active supplier, how a spend trend should be calculated, whether a quarter means calendar or fiscal, and when the agent should ask a clarifying question instead of assuming. These rules change with business practice, so they tend to live in a system prompt or a governed instruction file instead of the column metadata.

Anthropic recommends writing such instructions at the "right altitude": specific enough to guide behaviour, and general enough to give the model heuristics that apply to new questions (Anthropic). For a data agent, that means canonical definitions for the most contested terms, preferred analysis approaches and output formats, and guidance on when to ask.

Versions and Drafts

Every one of these parts will change. If the data model and instructions are edited in place, nobody can say which definition produced last month's number, and a bad edit reaches every user at once.

Treating context like code solves both problems. Versions give each answer a traceable definition, drafts let a team test a change before publishing it, and rollback limits the damage of a mistake. We return to this pattern in our guide on versioning AI agent context.

Should Context Live in the Prompt or the Data Model?

Teams often start by pasting everything into one long system prompt. That works for a demo, but long prompts dilute attention, a problem described in Context Rot: Why LLMs Degrade as Context Grows, and a prompt is a snapshot. When a new contract status appears in the data, the prompt stays the same until someone remembers to edit it.

A cleaner split assigns each kind of context to the place where it stays correct. Business practice goes in versioned instructions. The truth of the data (tables, keys, descriptions and current values) goes in the data model, where it can be refreshed from the source.

Approach

What it captures well

What it misses

How it stays current

Raw schema only

Table and column names, types

Meaning, source of truth, real values, metric rules

Automatically, but carries little meaning

Long system prompt

Business vocabulary, analysis approach, output format

Live values, keys, which table to use at scale

Only when someone edits it

Compiled semantic layer

Governed metrics with deterministic SQL

Questions outside the modelled metrics, judgement calls

Through a modelling workflow and deploys

Agent data model plus versioned instructions

Descriptions, keys, live values, and canonical definitions together

Rules nobody has written down yet

Values refresh on a schedule; definitions change by version

The last row is the pattern this guide recommends, with or without a semantic layer alongside it. Anthropic's guidance on "just in time" context points the same way: give the agent lightweight references and let it load details with tools when it needs them, instead of front-loading every fact (Anthropic). A data model is the index those tools read from.

How Do You Keep an Agent's Picture of the Data Fresh?

Data drifts even when nobody touches the agent. Operations adds a Suspended contract status, or a migration renames a payment term. Each change makes some part of a static prompt wrong, and the agent keeps filtering on values that no longer exist or ignores ones that now matter.

Three habits keep the picture current:

  1. Refresh values from the source. Mark the columns that users filter on and cache their distinct values on a schedule, so a new status appears in the agent's context the next morning without a manual edit.

  2. Version definitions, and review changes. When a definition changes, publish it as a new version with a note on what changed, so answers before and after the change can be explained.

  3. Test against known answers. Keep a set of questions with agreed results and rerun them when the data model changes. This is the subject of our guides on golden questions for AI agent evals and data drift in AI agents.

Over time, each gap the agent exposes becomes one more line in the data model, a loop described in Self-Improving Agents.

How Upsolve Approaches the AI Agent Data Model

Upsolve grounds every agent in two sources of context that are kept deliberately separate: a versioned system prompt for business practice, and an Upsolve data model for the truth of the underlying data. Neither replaces the other.

In our procurement demo project, a sample agent answers questions about spend trends. Its system prompt is on version 22, with a draft holding the latest unpublished changes. That prompt contains canonical definitions and vocabulary, analysis approaches, guidance on when to ask clarifying questions, and output formats. A team can change a definition, test the draft and roll back if something breaks.

The agent is attached to a data model, procurement data version 7, which registers the tables and fields the agent may use. Each table and column carries a description, a type, and its primary or foreign key role. Columns can be marked as selectable, which tells Upsolve to pre-cache their values. In the sample data, the agent knows the five contract statuses, the four payment terms and the four purchase-order statuses before it writes any SQL. Cached values refresh nightly by default, or on any schedule you set, so a new status in the data reaches the agent without anyone editing a prompt.

As Serguei Balanovich, CTO of Upsolve AI, puts it in the demo: "An agent has to be grounded in reality. Without that, you can have ten different definitions of ARR or monthly spend trend."

No semantic layer is required: you can build the data model directly on your warehouse tables and add definitions as you find gaps. The data model is also the base for the rest of the stack. It carries the data security rules covered in Part 02, Row-Level Security for AI Agents, and later layers in the series measure how well it holds up. For the surrounding tooling, see Context Engineering Tools.

Frequently Asked Questions

What is a data model for AI agents?

A data model for AI agents is a registered description of the data an agent can query, written for a language model to read. It includes table and column descriptions, data types, primary and foreign keys, and the real values of columns users tend to filter on. Combined with canonical metric definitions, it tells the agent which table is the source of truth, how tables join and what each value means. The agent then uses your organisation's definitions instead of inferring them from column names.

Is an AI agent data model the same as a semantic layer?

They overlap but are not identical. A semantic layer defines governed metrics, often compiled into SQL by a query planner, so every consumer computes them the same way. An agent data model describes every table and column the agent can see, including keys and live values, so the agent can handle questions no metric definition anticipated. Our semantic layer guide covers the metric side in depth.

Do I need dbt or an existing semantic layer to build one?

No. You can build a data model directly on top of the warehouse tables you have now. Start by registering the tables the agent should use, writing descriptions for the columns that cause confusion, recording keys, and marking the categorical columns users filter on. Add canonical definitions for the most contested metrics in versioned instructions. Existing dbt models are a useful source of descriptions and relationships, but they are not a prerequisite.

Should metric definitions go in the system prompt or the data model?

Put each kind of context where it stays correct. Business practice, such as how your team defines a spend trend or formats an answer, fits in versioned instructions because it changes when the business changes. Facts about the data, such as which table holds invoices, how tables join and which statuses exist, belong in the data model because they can be refreshed from the source. Keeping the two apart means a new status in the data does not require a prompt edit.

How often should an AI agent's data context be refreshed?

Refresh cached column values at least daily for most business data, and more often for columns that change within the day. Nightly refreshes catch new statuses, categories and terms before the next working day's questions. Definitions change less often and should be updated through a versioned review instead of a schedule. Questions that return zero rows are a useful warning that the agent's list of values is stale.

How do you stop an AI agent from using the wrong metric definition?

Write one canonical definition for each contested metric and make the agent read it before answering. Tell the agent which table is the source of truth, record keys so joins cannot double count, and give it the real values of filter columns. Add an instruction to ask a clarifying question when a term has more than one accepted meaning. Then rerun a set of questions with agreed answers each time the data model or definitions change.

Sources

  1. NeurIPS / Li et al.: Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs. https://arxiv.org/abs/2305.03111

  2. arXiv / Wretblad et al.: Synthetic SQL Column Descriptions and Their Impact on Text-to-SQL Performance. https://arxiv.org/abs/2408.04691

  3. arXiv / Talaei et al.: CHESS: Contextual Harnessing for Efficient SQL Synthesis. https://arxiv.org/abs/2405.16755

  4. arXiv / Gao et al.: Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation. https://arxiv.org/abs/2308.15363

  5. dbt Developer Hub: Join logic (MetricFlow). https://docs.getdbt.com/docs/build/join-logic

  6. Snowflake Documentation: YAML Specification for Semantic Views. https://docs.snowflake.com/user-guide/views-semantic/semantic-view-yaml-spec

  7. Anthropic Engineering: Effective context engineering for AI agents. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents

Give your agent one definition of every metric. Start free at upsolve.ai with 2,000 free AI credits, register your first data model, and see how your agent's answers change. Prefer a walkthrough? Talk to our team. Read more about Upsolve Data Models, and follow the rest of The Complete Agent Stack series.

Try Upsolve for Embedded Dashboards & AI Insights

Embed dashboards and AI insights directly into your product, with no heavy engineering required.

Fast setup

Built for SaaS products

30‑day free trial

See Upsolve in Action

Launch customizable dashboards and AI‑powered insights inside your app, fast and with minimal engineering effort. No code.

Follow us

Related Articles

Stop answering the same 10 questions today.

The Platform for Accurate, Reliable, and Trustworthy AI Analytics.

Agent Studio for Data Teams. Encode context. Deploy agents. Deliver clarity.

© 2026 Upsolve AI, Inc.  

|