Working with AI agents
Reports are plain code in a git repository, which is exactly the shape a coding assistant works well in: it can read what exists, write something new, and open a pull request you review. This page covers how to set that up so the first draft is usually right.
Write AGENTS.md first
Most bad output comes from missing context rather than a weak model. AGENTS.md
is where you put what a new colleague would need on day one:
# Working in this repository
## Data
- The warehouse is authoritative. Never query the replica.
- Revenue is defined in context/metrics.md. Do not derive it another way.
- `orders_legacy` is deprecated — use `orders`.
## Reports
- SQL lives in queries.py, never inline in generator.py.
- One report per directory under reports/.
- Every report declares its data sources in report.yaml.
## Review
- Open a pull request. Nothing is merged unstested.
- If a metric definition is unclear, ask rather than guess.
Be specific and be blunt. "Do not guess a definition — ask" prevents more bad reports than any amount of prompting.
Give it real data access
An assistant that can check a column writes SQL that runs. One that cannot will invent column names that look plausible and fail at build time — or worse, silently return the wrong grain.
Two ways to provide access, and you do not need both:
Configured data sources — the default
A source in data-sources/config.yaml is queryable directly, through the same connection a built report uses.
df = query("warehouse", "select ...")
No extra services. Works offline, in CI, in a notebook.
An MCP server — optional
For letting an assistant explore: browsing a schema it has never seen, or reaching a system that is not a configured source at all.
Complements the above; never a prerequisite for it.
Configured sources are usually enough — see Ad-hoc analysis.
Connect it to the portal
The portal is also an MCP server, so the agent that writes reports can read what the portal has built and, with the right key, operate it: check a report's last build and validation, query its published numbers, build it, test a data source, publish pending repository changes. The agent brings its own model — the portal holds no model key for this, and it works on a network with no route out.
Two things to have at hand:
- The server URL —
https://<portal>/s/<org>/<studio>/mcp: the studio's own URL plus/mcp, one server per studio. - A personal API key, sent as a bearer token. The agent acts as you, in that organization, under your studio role.
One command
With the key in hand, run this in the repository root — the API-keys page shows it ready-made, with the URL filled in, for each studio you can reach:
trellum setup portal --url https://<portal>/s/<org>/<studio> --key trellum_pk_…
It writes the two files below, puts the key in .env as TRELLUM_API_KEY,
and adds a short "Portal" section to CLAUDE.md and AGENTS.md so the
agent knows the server exists and what it is for. It refuses to finish while
.env is not ignored by git, and re-running it with a new URL or key
updates the entry in place without touching any other server in either
file. trellum doctor warns about an entry whose URL is malformed or whose
key is missing from .env. What it writes, if you would rather write it
yourself:
Claude Code
.mcp.json in the repository root:
{
"mcpServers": {
"trellum": {
"type": "http",
"url": "https://<portal>/s/<org>/<studio>/mcp",
"headers": { "Authorization": "Bearer ${TRELLUM_API_KEY}" }
}
}
}
Cursor
.cursor/mcp.json, the same shape:
{
"mcpServers": {
"trellum": {
"url": "https://<portal>/s/<org>/<studio>/mcp",
"headers": { "Authorization": "Bearer ${env:TRELLUM_API_KEY}" }
}
}
}
Both files reference the key through an environment variable and are safe to
commit. The key itself lives in your shell environment or the .env the
repository already keeps out of git — never in the file.
read or write
The key's scope decides what the agent is offered:
| Scope | The agent gets |
|---|---|
read |
The four read tools. It can look at anything you can look at and change nothing. |
read,write |
The read tools plus the actions — and they run immediately. There is no approval card here: your coding agent asks you before every tool call, and that is the approval. |
Actions are offered only when an organization admin has turned on Let the assistant propose actions (the same switch the in-portal assistant uses) and your studio role allows them; a viewer gets the read tools whatever the key's scope. Every action is recorded in the audit log against you, with the key's id.
The tools
| Tool | What it does | Needs |
|---|---|---|
list_reports |
The studio's reports: slug, name, category, last build, link | read |
get_report_details |
One report's charts, datasets and columns, validation state, last run | read |
query_report_data |
Filter and aggregate a report's already-built data — published numbers, no warehouse query | read |
read_doc |
A report's report.yaml or queries.py, assistant.md, project docs (allowlisted paths) |
read |
check_repo_changes |
What the portal would publish: the remote head it last saw and when, reports and data-source declarations added, changed or removed since the last publish, the last error | read, developer |
test_data_source |
Re-run a data source's connection test | write, developer |
run_report |
Queue a build now | write, developer |
publish_repo_changes |
Publish fetched-but-unapplied repository changes (manual publish mode) | write, developer |
configure_data_source |
Point you at the page where a declared source's credentials are entered — see below | write, studio admin |
Secrets never travel over MCP
configure_data_source takes no password, key or token. It answers with the
link to the studio's Configure page for that source and the fields it still
needs; you type them there, in the browser, and they go straight to the
encrypted store. A secret passed as a tool argument is refused and nothing is
stored — so a credential never lands in your agent's context or transcript.
Publish and configure from your agent
Once wired, the agent that edits reports can also take them to the portal
and back. trellum guide portal prints this loop to the agent; here it is
for you.
- Edit and build locally.
python -m trellum.run reports/<slug> --no-serve. - Validate.
python -m trellum validate reports/<slug>— a report is finished when there are no FAILs. - Commit and push. The portal reads the branch it is configured for, never a working tree.
- Check.
check_repo_changessays what the portal saw at its last fetch: the remote head and when, the reports added, modified and removed since the last publish, data-source declarations that changed, and the last error. In manual publish mode this is the review step; in auto mode the fetch already published. - Publish.
publish_repo_changespublishes the pending changes at the head the agent just reviewed. - Build and read the result.
run_reportqueues a build; a moment laterget_report_detailsshows the last run, its status and the validator summary the portal recorded. A FAIL here is the same FAIL the local validator shows: the fix is in the repository. - Fix a data source. A build that fails in the driver usually means a
credential:
test_data_sourcere-runs the connection test and reports the detail line;configure_data_sourcereturns the studio's Configure link, where you enter the secret in the browser. The agent never sees it. - Add an alert.
create_alertturns "tell me if APAC DAU drops sharply" into a rule on a report — evaluated after each build or on a schedule by the portal's own agent, see Alerts;update_alertchanges it.
Steps 4 and 6 read, and work on a read key. Steps 5, 7 and 8 need
read,write and the organization's actions switch; your harness asks before
each call, and every call is in the audit log under your key.
A good first prompt
Point at the context explicitly rather than assuming it will be found:
Read
AGENTS.mdandcontext/metrics.md. Add a report underreports/churn-cohorts/showing monthly churn by signup cohort for the last twelve months, using the churn definition in the metrics file. Follow the structure ofreports/weekly-revenue/. Runpython -m trellum.run reports/churn-cohorts --no-serveand fix anything that fails before you finish.
Naming an existing report as the pattern to follow is the single highest-value part of that prompt.
What to check before merging
The diff is the review surface, so review the diff:
- Does the SQL match the definition? The most common failure is code that runs perfectly and computes the wrong thing.
- Is the grain right? A join that duplicates rows inflates a total without any error appearing.
- Are filters where they belong? A
whereclause in the wrong place quietly changes what a number means. - Did it invent a column? If the build succeeded, it did not — which is why it should build before you review.
- Is it in
queries.py? Inline SQL is harder to review and harder to reuse.
Warning
An assistant is fast at producing plausible analysis. The reviewer supplies the judgement about whether it is correct — that responsibility does not move. The workflow is built so review is possible: a diff you can read rather than a chart you have to trust.
Why this beats asking a chatbot for a number
A chat answer is produced once and cannot be re-derived — you cannot audit it, re-run it next quarter, or explain in a meeting how the number was reached. Here the assistant's output is a program: reviewed, committed, and re-run on a schedule. Everyone who opens the dashboard next quarter sees a number produced by that same approved code.