# README

Dot, the data bot is your AI data analyst.

<figure><picture><source srcset="/files/GKXZckYKXCNZ1a6kmtyC" media="(prefers-color-scheme: dark)"><img src="/files/7fpXlQ5UOZdlNmcIcgo6" alt=""></picture><figcaption></figcaption></figure>

Answer most business questions instantly 24/7. Data teams can focus on deep work, not on answering easy questions about dashboards.

Dot, the data bot is an intelligent virtual data assistance that answers business data questions, retrieves definitions and relevant data assets, and can even assist with data modelling.

Leverages the power of large language models combined with your data stack (Snowflake, BigQuery, Redshift, see all [Integrations](/integrations/databases)) and documentation of processes and metrics.

Ships with a developer training space that ensures the data team is in the driver seat and results from Dot are validated. Say No to hallucinations.

Ready for Enterprise usage with SOC 2 Type II compliance. Your data is a treasure. We keep it safe.


# Getting Started

Get answers from your data in under 10 minutes.

Not using Dot yet? [Sign up for free.](https://app.getdot.ai/register)

***

## 1. Connect Your Data

Go to **Settings → Connections** and add your data warehouse, semantic layer, or BI tool.

No code required—just permissions.

<figure><img src="/files/RnAUVSo1P4fE7jBExfKd" alt="Connect databases like Snowflake, BigQuery, Postgres, or semantic layers like dbt, Looker, and Power BI"><figcaption></figcaption></figure>

[See all integrations →](/integrations/databases)

***

## 2. Describe What You Need

Open the **Context Agent** and tell it what you're trying to accomplish. It will interview you, explore your data, and set things up.

<figure><img src="/files/y4tAoSl3YZOaLA50znDM" alt="Context Agent interviewing a user about their product analytics use case"><figcaption></figcaption></figure>

You don't need to configure everything upfront—start with one use case and expand from there.

[Learn more about the Context Agent →](/train-dot/context-agent)

***

## 3. Start Asking Questions

Ask in plain English. Dot queries your data, visualizes results, and explains what it found.

You can ask questions directly in Dot, or connect Slack or Microsoft Teams so your whole team can get answers where they already work.

<figure><img src="/files/ulpi96XgUdEguRGYec2k" alt="Team members asking Dot questions in Slack and getting answers with charts"><figcaption></figcaption></figure>

[Set up Slack & Teams →](/integrations/slack-and-teams)

***

## What's Next

**Fine-tune how Dot understands my data**\
[Model →](/train-dot/model)

**Control who can access what**\
[Permissions →](/train-dot/permissions)

**Schedule recurring reports**\
[Scheduling →](/using-dot/scheduling)


# Analyze

Ask anything about your data in plain English. Dot investigates, writes SQL, and answers with cited, verifiable results.

Your data lives in complex models, and the answers you need are scattered across hundreds of dashboards. The few people who can navigate all of it become a bottleneck.

Dot removes the bottleneck. Ask a question in your own words and Dot does the analysis: it understands your data model, writes and runs the SQL, checks its own work, and answers with a clear result you can trace back to the source. No ticket, no dashboard hunt, no waiting on the data team.

<figure><img src="/files/EBLq1t4oKtEMQn7UCcsR" alt="Dot answering a revenue-growth question with a cited headline and a rising line chart"><figcaption><p>Ask in plain English → Dot investigates and answers, with every number traceable to its source.</p></figcaption></figure>

## One place for every question

There's no mode to pick. A quick lookup and a deep investigation start the same way — you just ask. Dot reads the question and does as much work as it needs:

* **"What was revenue last month?"** → a direct answer with the number, a chart, and the query behind it.
* **"Why did revenue drop in Q3?"** → an autonomous investigation. Dot runs multiple queries, tests explanations, rules out dead ends, and comes back with a structured report — headline, key findings, supporting charts, and recommended next steps.

Under the hood Dot can understand your semantic model, write and run SQL, compute and reshape data, build visualizations, and search the web for context when a question needs it. The harder the question, the more of these it chains together on its own.

## Energy modes

One dial controls how much thinking Dot puts into an answer. Open the **energy mode** selector next to the message box:

* **Economy** — lighter on cost, a solid pick for everyday questions.
* **Balanced** — the default. More depth when a question needs it.
* **Frontier** — the deepest reasoning, for your most complex questions.

<figure><img src="/files/gcLvl8OGhj6X4RM5V7TT" alt="The energy mode selector showing Economy, Balanced, and Frontier"><figcaption><p>Pick the effort to match the question — Economy for quick looks, Frontier for the hard ones.</p></figcaption></figure>

Your choice sticks to the conversation, so a follow-up keeps the same depth. Admins can set the default for the whole workspace and, if they prefer, lock it (see [Configuration](#configuration-admins)). When it's locked, everyone uses the workspace default, and your own pick, including the Slack and Teams keywords below, is ignored.

{% hint style="info" %}
In Slack, Teams, and email you can set the mode with a keyword: start a message with `!economy`, `!balanced`, or `!frontier`. Adding **"go deep"** or **"ultrathink"** to any message bumps just that one to Frontier.
{% endhint %}

## Answers you can trust

Dot is grounded in your connected data, so it doesn't invent numbers — every answer is backed by a query you can inspect.

* **Trace any number.** Each answer carries **Chart / Data / Query / Code** views and a **Source** pill. Click through to see the exact SQL and the rows it ran on.
* **Cited claims.** Statements in a report link back to the data that supports them, so anyone can verify a finding instead of taking it on faith.
* **A full investigation log.** For multi-step work, open **Full logs** to see every query and step Dot took to reach its conclusion.

If an answer looks off, just say **"try again."** Dot is built on frontier models and is statistical by nature — a second pass often takes a different, better route. And if you're ever unsure, the data team can audit the query and the model behind any answer.

## Keep the conversation going

Dot holds context across a thread, so you refine instead of restarting:

* *"Drill into EMEA."*
* *"Show the top 10 customers driving this."*
* *"Now compare it to last year."*

You can also attach a file (like a CSV) to bring your own data into the analysis.

## Do something with the answer

Every answer is a starting point, not a dead end:

* **Share** a link with a colleague, or **schedule** it to run and deliver on a cadence.
* **Download as a presentation** to drop findings straight into a deck.
* **Turn it into an** [**App**](/using-dot/apps) — a living, interactive dashboard or report that stays connected to your data.

## How to get the best answers

Dot is good at figuring out intent, but a clear question gets a sharper answer. The more of these you include, the less Dot has to guess:

* **The metric** — revenue, active buyers, churn rate.
* **The entity** — customers, products, suppliers.
* **The time frame** — last month, Q1 2026, year to date.
* **The grain and cuts** — by month, by region, for Enterprise only.

So instead of *"How are sales?"* ask *"What was Enterprise revenue by month for the first half of 2026, in euros?"* And when you're not sure what's available or how to phrase something, just ask Dot for help — it knows your data model and will suggest a direction.

## Configuration (Admins)

In **Settings → Advanced Settings → AI & Data Configuration**:

* **Default Energy Mode** — set the starting mode for the workspace, and optionally lock it so everyone uses the same depth.
* **Business questions only** — off by default. See below.

On the Model page, under **Design**:

* **Custom appendix** — append a disclaimer or contact details to every answer.

Deeper questions do more work, which uses more [Agent Compute Credits](/train-dot/agent-compute-credits). Admins can cap usage per user or group to keep spend predictable.

### Business questions only

If people use Dot for things it is not there for, you can keep it on topic. Turn on **Business questions only** and Dot answers work questions and politely declines the rest, like weather, jokes, sports, or homework. It says so in two sentences, in the person's own language, and names something it can help with instead.

Work questions are not limited to your connected data. Competitors, suppliers, customers, market size, pricing, regulation, and industry benchmarks all count, and Dot researches them on the web as usual.

This is useful for spend as well as for focus. Dot decides before it does anything, so a declined question runs no query and no search, and costs close to nothing.

{% hint style="info" %}
The setting applies to the workspace you turn it on in. Turn it on separately in each workspace you want it in. It guides Dot rather than blocking it, so treat it as a strong steer, not a hard filter.
{% endhint %}


# Build apps

Turn any analysis into a living, shareable data app — dashboards, reports, and presentations that stay connected to your data.

A chat answer is a snapshot in time. A Dot **app** turns that analysis into a **living data product**: you ask for it in plain English, Dot builds it, and it stays connected to your data — fast, interactive, shareable, and traceable to the source.

Apps sit between hand-configuring a BI dashboard and coding one from scratch. You describe what you want; Dot writes the queries, compiles them to real SQL once, and pins the result. Every view then re-runs the real query against live data — **no AI in the render path**, so it's fast and deterministic — while everything stays editable in plain English.

<figure><img src="/files/rfALLwXV9X8X3RNhuEZ8" alt="Asking Dot to build an executive overview dashboard, with the live app rendered alongside the chat"><figcaption><p>Ask in plain English → Dot builds the app → publish, share, schedule.</p></figcaption></figure>

## What you can build

Apps are one format with several modes — **dashboards are just the most common one**:

* **Dashboards** — track KPIs and trends: filterable grids of metrics, charts, and tables.
* **Reports** — written, editorial analyses ("state of the business", post-mortems, field notes) where every number is cited.
* **Presentations** — live web slide decks for a QBR or board review, with a **Present** button.
* **Data essays** — scrollytelling narratives where a single graphic evolves as you scroll.

They all share the same engine, primitives, and data connections — the difference is a layout choice, not a different tool.

<figure><img src="/files/tvRlDlQnF458gGE7PWlN" alt="An editorial data report titled Product usage field notes and trends"><figcaption><p>A <strong>report</strong> — the same data, as a written, cited analysis.</p></figcaption></figure>

<figure><img src="/files/wtGQwMHcAp6mKOII9ZQS" alt="A board-review presentation slide titled Growth is expansion-led, with a monthly MRR bridge chart"><figcaption><p>A <strong>presentation</strong> — a live board deck you can present full-screen.</p></figcaption></figure>

## How you make one

1. **Ask in chat** — "build an executive overview dashboard", "turn this into a board deck". (Or click **New App**.)
2. **Dot builds it** — it compiles your plain-English questions to SQL, dry-runs them, and shows a live preview with an **Open app** card.
3. **Refine conversationally** — "add a region filter", "make it a report", "change the theme". Every edit rebuilds and re-previews.
4. **Publish** — the app appears on your **Apps** page, either personal or shared with the workspace.

## Live, interactive, and yours to edit

* **Always current** — every view runs the pinned SQL against live data. Refresh on demand or on a [schedule](/using-dot/scheduling).
* **Interactive** — a declarative filter threads through every card whose data has that column and cross-filters the rest automatically; view controls (log/linear, time window, smoothing, show/hide series) flip instantly, client-side, with no re-query. Your filter choices go into the address bar, so copying the URL shares the exact view you are looking at.
* **Editable** — AI-built, but fully human-editable: the text, layout, style, and the queries themselves — in chat or by hand.

## Share it anywhere

* **Share link** — a private link you can revoke at any time.
* **Embed** — drop an app into your wiki, portal, or product.
* **PDF** — a board-ready export that still links back to the live app. It exports the view you are looking at, so the filters you picked are the ones in the file, and the link inside it reopens that same filtered view.
* **Present** — full-screen live slides straight from the browser.

## Trust and governance

* **Every number traces to its source.** Click any value to open the exact query and its compiled SQL; each card carries a **Source** pill, and the data-lineage view maps an app back through its queries, tables, dbt models, and sources.
* **Certification.** An admin or modeler can mark an app **Certified** — and the badge drops automatically if the underlying source changes without re-review, so a trust signal never goes stale silently.
* **Permissions & usage.** Folder-based view/edit control, per-app view counts, and auto-archiving of apps no one opens anymore.

## Apps as code

For teams who review changes like software, every app is a plain `.app` file you can version-control. Pull and edit it locally, push to recompile the SQL, and open a pull request — with model changes branched and merged through environments.

See [CLI & AI Agent Skill](/developers/cli) for `dot apps` and `dot env`, and [GitHub Sync](/train-dot/version-control/github) for keeping apps in your repo.

{% hint style="info" %}
"Dashboard" is one kind of app. The feature grew from dashboards → reports → **apps**, which is now the umbrella for all of them.
{% endhint %}


# Scheduling

Automate recurring reports

Schedule Deep Analysis reports to be delivered automatically via Email, Slack, or Teams.

Unlike dashboard snapshots that show numbers, scheduled reports explain what changed and why—with trends, anomalies, and recommendations included.

### Use Cases

**Meeting prep**: Schedule 30 minutes before recurring meetings. Your team gets context without manually checking dashboards.

**Weekly reviews**: "What changed this week?" delivered Monday morning to Slack—ready for discussion.

**Replace dashboard check-ins**: Instead of pulling data, insights come to you.

### Creating a Schedule

1. Run a Deep Analysis query
2. Click the **Schedule** button on the response

<figure><img src="/files/rDN2YczqOZHKUp0g5y3H" alt=""><figcaption><p>Click Schedule to set up recurring delivery</p></figcaption></figure>

3. Choose delivery channel (Email/Slack/Teams)
4. Add recipients
5. Set frequency: Daily, Weekly (pick day), or Monthly (pick date)
6. Click **Schedule**

<figure><img src="/files/KSKgpsaVGzyPprBJwoHm" alt=""><figcaption><p>Configure channel, recipients, and frequency</p></figcaption></figure>

Click **Test** to send a preview before committing.

### Managing Schedules

Open the schedule modal on any scheduled message to:

* Edit frequency or recipients
* View run history (past deliveries and costs)
* Delete the schedule

### Work Gate — Skip Runs When There's Nothing New

Avoid wasting credits on scheduled runs when conditions aren't met. A work gate checks your database before the agent runs — if the check fails, the run is skipped entirely and no credits are consumed.

**Use cases:**

* Skip a daily sales report if no new orders arrived
* Don't run a pipeline health check until the ETL job has finished
* Only analyze data after a specific table was updated

Dot writes and maintains the gate script for you based on a condition you describe. You can test the gate against live data before activating it. Skipped runs show as `work_gate_blocked` in run history.

{% hint style="info" %}
If the gate encounters an error, the agent runs anyway (fail-open) — so you never silently miss a report.
{% endhint %}

***

### Result Gate — Only Deliver When It Matters

Reduce notification noise by setting a plain-English condition that controls whether the report is actually sent to recipients.

The agent still runs the full analysis, but only delivers the report if your condition is met. No extra cost — the condition is evaluated as part of the normal analysis.

**Examples:**

* *"Only send if monthly revenue decreased by more than 5%"*
* *"Send only when there are more than 3 anomalies detected"*
* *"Deliver if customer churn rate exceeds the previous month"*

Suppressed reports show as `result_gate_suppressed` in run history, so you can always review what was analyzed but not delivered.

***

### Combining Both Gates

You can use work gate and result gate together on the same schedule:

1. **Work gate** checks if there's reason to run at all (saves credits)
2. **Result gate** checks if the findings are worth delivering (reduces noise)

Example: A daily pipeline report with a work gate that skips if the ETL hasn't finished, and a result gate that only sends when the error count is above threshold.

***

### Full analysis delivery

By default a scheduled report sends a short executive summary. If you'd rather send the whole thing, with all the charts, tables, and files, turn on "Include full analysis" when you set up the schedule.

This is a per-schedule choice. You can send the full report for one schedule and keep another as a short summary. The toggle sits in the schedule dialog, next to where you pick the channel and recipients, and you can change it later by opening the schedule again.

Here's what each channel sends either way:

| Channel | Summary (default) | Full analysis                    |
| ------- | ----------------- | -------------------------------- |
| Email   | Summary PDF       | Full report with charts and data |
| Slack   | Summary message   | Complete analysis with charts    |
| Teams   | Summary card      | Full card with charts and a PDF  |

***

### Costs & Limits

* **1 ACC** (Agent Compute Credit) per scheduled run
* **Maximum frequency**: 2 times per day
* Work gate blocks do **not** consume an ACC
* Result gate suppressions **do** consume an ACC (the agent ran, only delivery was gated)

### Admin Feature: Run As User

Admins can schedule reports to run as another user—useful for delivering reports with that user's data permissions.


# Root: Context Agent

Your AI data team member that builds and maintains your knowledge base

Root is an AI agent that helps you build, maintain, and evolve your organization's knowledge base. It runs in an isolated sandbox with access to your connected tools—databases, BI dashboards, and past conversations—and can create documentation automatically.

<figure><img src="/files/y8qwAsabeGoT6GTEyhaU" alt=""><figcaption><p>Root helps you curate and share company knowledge with Dot</p></figcaption></figure>

**Why this matters**: Building a comprehensive knowledge base manually takes months. Root accelerates this by extracting business logic from your existing systems and learning from how your team actually uses data.

Root also learns on its own. When Dot spots something worth remembering during a conversation, like a corrected metric, a renamed table, or a missing definition, it sends a proposal to Root. Someone reviews it and merges it with one click if it looks right. That's an admin by default, and you can let modelers review proposals too. Dot gets smarter every day, and people stay in control.

***

## Workflow

1. **Open Context Agent** from the sidebar
2. **Describe what you need** in natural language—Root understands complex requests
3. **Approve tool use** when Root needs to query data or make changes

<figure><img src="/files/CBAGLdkCdChODw1kPKES" alt=""><figcaption><p>Review and approve changes before they're applied</p></figcaption></figure>

4. **Review the diff** to see exactly what changed

<figure><img src="/files/DC6btlvEUYTR0lyfPEhh" alt=""><figcaption><p>Click Review Changes to see pending modifications</p></figcaption></figure>

5. **Merge to production** when satisfied—or discard and try again

<figure><img src="/files/86G7lrntKxwuKcQUU0QS" alt=""><figcaption><p>Review the diff and merge when ready</p></figcaption></figure>

All changes happen in an isolated sandbox. Nothing goes live until you explicitly merge.

***

## Use Cases

### 1. Extract Metrics from BI Tools

**Problem**: Your Tableau/Metabase dashboards contain business logic, but it's not documented anywhere Dot can use.

**Solution**: Give Root access to your most trusted dashboards and ask it to create a metric glossary.

```
Here are our five most trusted Metabase dashboards—analyze them and
create a glossary of key metrics with their definitions.
```

Root will:

* Connect to your BI tool via API
* Extract calculations, filters, and business logic
* Create standardized metric definitions Dot can use

***

### 2. Learn from Past Conversations

**Problem**: You don't know what questions your team asks most or what's missing from your documentation.

**Solution**: Ask Root to analyze past Dot conversations.

```
Analyze the last 30 days of conversations. What are the most
common questions? Are there patterns in failed queries?
```

Root will:

* Export and analyze conversation history
* Identify frequently asked questions
* Find gaps where Dot couldn't answer
* Suggest documentation improvements

***

### 3. Audit Existing Documentation

**Problem**: Your table descriptions were written months ago. Are they still accurate?

**Solution**: Ask Root to find inconsistencies.

```
Are there inconsistencies in our documentation or data source
descriptions? Check if sample values match descriptions.
```

<figure><img src="/files/lmllLcpIfF3xwQ2n9yEK" alt=""><figcaption><p>Root analyzes your context and identifies inconsistencies</p></figcaption></figure>

Root will:

* Read your current documentation
* Query actual data to verify descriptions
* Flag mismatches between docs and reality
* Suggest fixes

***

### 4. Interview-Based Knowledge Capture

**Problem**: Tribal knowledge exists in people's heads, not in documentation.

**Solution**: Let Root interview domain experts and capture their knowledge.

```
Interview me about how we handle sales and create a note.
```

Root will:

* Ask targeted questions about your process
* Capture answers in structured notes
* Create documentation that reflects actual practice

***

### 5. Bulk Table Documentation

**Problem**: You have hundreds of tables but no descriptions.

**Solution**: Point Root at your schema and let it document everything.

```
Activate all tables in the 'REPORTING' schema. Add descriptions
based on column names and sample values.
```

Root will:

* Query database metadata
* Analyze column names, types, and sample data
* Generate descriptions for each table and column
* Save as documentation Dot can use

***

### 6. Migrate Documentation

**Problem**: Your documentation lives in Confluence/Notion, not where Dot can use it.

**Solution**: Ask Root to migrate it.

```
Here's a link to our Confluence space. Extract the key business
definitions and create notes for Dot.
```

Root will:

* Fetch content from external sources
* Extract relevant business context
* Create notes in Dot's format

***

### 7. "Remember This"

**Problem**: Someone on your team knows that "fiscal year starts in April" or that the `orders` table was renamed to `transactions` last month — but that knowledge is stuck in their head.

**Solution**: During any conversation, just tell Dot to remember it.

```
Remember that our fiscal year starts in April, not January.
```

Dot will propose a knowledge base update. A reviewer sees the proposal, checks what changed, and merges or rejects it. That's an admin, or a modeler you've given the review permission to.

The people closest to the data are the ones who catch mistakes first. This lets them fix things on the spot, with an admin verifying before it goes live.

***

### 8. Investigate a Chat

**Problem**: A user had a bad experience — wrong numbers, a confusing chart, or a query that missed the point. You want to know why.

**Solution**: Open the chat from your history and ask Root to look into it. Root reads everything Dot did during that conversation, which tables it picked, which SQL it wrote, and where it went wrong, and explains the root cause.

```
Look into this chat and make sure this mistake doesn't happen again.
```

Root will:

* Trace every decision Dot made in that conversation
* Identify where things went wrong
* Propose a fix — a corrected note, a missing relationship, or a clearer description

Instead of manually debugging, you get a diagnosis and a fix in one step.

***

### 9. Find Recurring Issues

**Problem**: The same type of mistake keeps happening across different users and conversations, but nobody has connected the dots.

**Solution**: Ask Root to look for patterns.

```
What are the most common errors across last week's conversations?
Suggest fixes for the top three.
```

Root will:

* Scan recent conversations for recurring failures
* Group them by root cause
* Propose targeted knowledge base improvements for each

This turns reactive troubleshooting into proactive improvement.

***

## How Dot Learns

Dot improves its knowledge base continuously, but nothing changes without a review:

1. **Dot spots something** — during a conversation Dot notices a mismatch, or someone says "remember this"
2. **A proposal appears** — proposals collect in the Proposals inbox, each with a clear diff of what would change
3. **A reviewer checks it** — an admin, or a modeler with the review permission, opens the proposal and sees exactly what's being added or corrected, and why
4. **Merge or reject** — one click to approve, or reject if it's not right
5. **Dot is smarter** — the improvement is live right away for everyone

Dot suggests. People decide.

***

## How It Works

1. **Start a session** from the sidebar (Context Agent)
2. **Ask Root** what you need—it understands natural language
3. **Review changes** before they go live (git-based versioning)
4. **Merge to production** when you're satisfied

All changes are version-controlled. You can pause, resume, or discard work at any time.

***

## What Root Can Access

| Source                  | Capability                                    |
| ----------------------- | --------------------------------------------- |
| **Databases**           | Execute SELECT queries, analyze structure     |
| **BI Tools**            | Read Tableau/Metabase dashboards via API      |
| **Past Conversations**  | Analyze Dot usage patterns                    |
| **Conversation Traces** | Replay and diagnose any past Dot conversation |
| **Web**                 | Search for documentation and best practices   |
| **Your Notes**          | Read and edit existing documentation          |

***

## Tips

* **Start specific**: "Document the orders table" works better than "document everything"
* **Iterate**: Root can refine its work—ask for changes if the first draft isn't right
* **Review diffs**: Always review changes before merging to production
* **Use interviews**: For complex processes, let Root interview you rather than trying to explain everything upfront
* **Review proposals regularly**: Dot learns fastest when proposals are reviewed quickly
* **Encourage "remember this"**: The people closest to the data catch the best corrections — let them contribute
* **Investigate disliked chats**: The fastest way to improve Dot is to diagnose what went wrong and fix it at the source


# Model

Train Dot on your data model, semantic layer, documentation, ...

Every good analysis is not just based on writing SQL, but on understanding how your business relates to the data you capture. The Model/Training space is for the data team to configure:

* Which data should Dot have access to?
* What does each table, column and row represent?
* What are existing analyses that can be build upon?

{% hint style="info" %}
**Edit safely with environments.** Every change here — selecting data, writing descriptions, defining relationships — can be made inside an [Environment](/train-dot/environments): an isolated copy of the model where you preview changes against real questions, then merge to production when you're ready.
{% endhint %}

## Select Data

**Datasource**\
Click on all the tables or explores that Dot will search through.

<figure><img src="/files/88BiNPIzdv1Ua0z0A3Y3" alt="" width="257"><figcaption></figcaption></figure>

**Fields**

Select all fields in a datasource that Dot should know about.

<figure><img src="/files/pks8sGHk5PmfWe886pmc" alt=""><figcaption></figcaption></figure>

## **Describe Data**

You can click suggest to automatically generate documentation and fetch sample values. Make sure to save your changes.

## Define Relationships / Joins

Joining data across different tables is really powerful because it allows us to answer much more broad questions about our business. However, joining data correctly is requires a good understanding of the relationships between your tables.

To ensure Dot avoids join fan-outs or the chasm trap, you can predefine relationships.

For each table, you would specify the foreign key references.

Note:

* a foreign key can be composed key, consisting of multiple columns
* foreign keys should usually refer to primary or natural keys

**Example**

<figure><img src="/files/AmXhxExp5qlwLyvfg2XS" alt=""><figcaption></figcaption></figure>

Given the data model above, you want to define 2 Relationships:

* For Orders
  * Foreign Key: `user_id`
  * Referenced Table: `Users`
  * Referenced Keys: `id`
* For OrderItems
  * Foreign Key: `order_id`
  * Referenced Table: `Orders`
  * Referenced Keys: `id`


# Notes

Context is all you need

Notes allow you to add useful context to Dot or to provide instructions on how to behave in certain situations. Helping business users with their data questions or creating an impactful churn analysis requires Dot to understand more about your business and your information architecture than can be found on the internet (usually).

{% hint style="info" %}
**Edit safely with environments.** Notes can be created and edited inside an [Environment](/train-dot/environments) — an isolated copy of the model — so you can preview how they change Dot's answers before merging to production.
{% endhint %}

Here are four types of notes that we found to be most useful:

## Organization Notes

This is Wikipedia style knowledge about your company, the products and services your selling, the internal processes that you follow, etc. While these things can change aim to go top down and start with the things that are probably still true two years from now.

**Example**

```
TinyTrucks is a children's toy company. 
We manufature educational and entertaining toy trucks.
We sell directly them online to parents in Europe and the US.
```

The example is a lot shorter than what you usually want to add as background.

## Agent Operating Principles

These are instructions on how to behave in certain situations and what principles to follow. The key is to write these as instructions and to use strong language with words like: `Always`, `Never`, `Only`

**Example**

```
Always clarify with the user if they want live or booked revenue when they ask about revenue.
Always by default filter out internal users (IS_INTERNAL = FALSE).
Always use the table fct_arr_consolidated when this table is enough to answer questions about ARR.
Never answer questions about a/b tests and rather refer them to https://experiments.internal.com
```

## Playbooks / Use Cases

For your high value use cases, you already have a lot of intuition and tribal knowledge on how to analyze certain situations. For example, when you review a marketing campaign, you have your go to data sources, you know which metrics you care most about, how you think about attribution etc.

Playbooks are about encoding and documenting this knowledge, so that the analytics agent can replicate it.

**Example**

{% code overflow="wrap" expandable="true" %}

````
# Use Case: Marketing Campaign Review Report

This prompt generates a comprehensive performance review for a digital marketing campaign, focusing on key channels like **Google Ads and Facebook Ads**. The report MUST analyze performance against KPI targets, segmented by **Channel, Audience Type (Prospecting vs. Retargeting), and Creative Format**. It provides a clear narrative on what's working and what isn't, projects end-of-campaign outcomes, and offers actionable recommendations for optimization.

---
## Important Notes:
- When referring to KPI performance, always specify which of the following you are referring to: **Spend-to-Date (STD), Campaign Target, Current Attainment, or Pacing Projection**.
- KPIs like **Cost Per Acquisition (CPA)** and **Cost Per Click (CPC)** are efficiency metrics; a value higher than the target indicates underperformance.
- **Return On Ad Spend (ROAS)** = (Total Conversion Value / Total Spend).
- All "overall" language MUST read **“Overall Campaign ROAS”** or **“Overall Campaign CPA”**.
- **Channel Groups**: Group all Google Ads campaigns → **Paid Search**; group all Facebook/Instagram Ads campaigns → **Paid Social**.
- **Audience Types**: Group campaigns targeting new users (e.g., based on interests, lookalikes, keywords) → **Prospecting**. Group campaigns targeting past website visitors or customer lists → **Retargeting**.
- Use **UTM parameters** as the primary source for consistent cross-channel tracking.
- Exclude all data where `utm_campaign` contains "test" or "internal".
- Ensure the **same aggregated dataset** is used for the summary narrative and the charts to avoid discrepancies.
- **Formatting**: Money as **$X,XXX** or **$X.XK**; percentages with **one decimal place + "%"**; ROAS as a ratio like **X.X:1**.
- **Chart Order**: In channel comparisons, always show **Paid Search, Paid Social**. In audience comparisons, show **Prospecting, Retargeting**.
- Use safe division (`NULLIF(denominator, 0)`) in all calculations.
---
## **Introduction**
Evaluating the performance of digital advertising campaigns requires a consolidated view across multiple platforms. Siloed reports from Google and Facebook make it difficult to assess overall effectiveness and make strategic budget decisions. This report template unifies critical data into a single, actionable review, allowing marketing teams to quickly identify top-performing channels, audiences, and creatives, and to optimize spend for maximum return.
---
## Definitions
### 1) Key Performance Indicators (KPIs)
- **Spend**: The total amount of money spent on advertising.
- **Impressions**: The number of times ads were shown.
- **Clicks**: The number of clicks on ads.
- **Click-Through Rate (CTR)**: The percentage of impressions that resulted in a click (`Clicks / Impressions`).
- **Cost Per Click (CPC)**: The average cost for each click (`Spend / Clicks`).
- **Conversions**: The number of desired actions completed (e.g., purchases, leads).
- **Cost Per Acquisition (CPA)**: The average cost for each conversion (`Spend / Conversions`).
- **Return On Ad Spend (ROAS)**: The total revenue generated for every dollar spent (`Revenue / Spend`).
### 2) Core Segments
- **Channels**: Paid Search (Google Ads), Paid Social (Facebook/Meta Ads).
- **Audience Types**: Prospecting (reaching new customers), Retargeting (re-engaging past visitors).
- **Creative Formats**: Video, Image, Carousel, Text Ad.
---
## Query Templates:
```
-- Fetch overall campaign KPIs vs targets
SELECT
  SUM(spend) AS total_spend,
  SUM(revenue) AS total_revenue,
  SAFE_DIVIDE(SUM(revenue), SUM(spend)) AS overall_roas,
  SAFE_DIVIDE(SUM(spend), SUM(conversions)) AS overall_cpa
FROM
  `your_project.your_dataset.marketing_campaign_data`
WHERE
  date BETWEEN '{{start_date}}' AND '{{end_date}}'
  AND campaign_name = '{{campaign_name}}';
-- Note: Targets are often stored in a separate table or spreadsheet and joined.
```
```
-- Breakdown performance by channel and audience type
SELECT
  channel,          -- 'Paid Search', 'Paid Social'
  audience_type,    -- 'Prospecting', 'Retargeting'
  SUM(spend) AS spend,
  SUM(revenue) AS revenue,
  SAFE_DIVIDE(SUM(revenue), SUM(spend)) AS roas,
  SAFE_DIVIDE(SUM(spend), SUM(conversions)) AS cpa
FROM
  `your_project.your_dataset.marketing_campaign_data`
WHERE
  date BETWEEN '{{start_date}}' AND '{{end_date}}'
  AND campaign_name = '{{campaign_name}}'
GROUP BY
  channel, audience_type;
```
```
-- Get daily ROAS to plot performance over time
SELECT
  date,
  SAFE_DIVIDE(SUM(revenue), SUM(spend)) AS daily_roas
FROM
  `your_project.your_dataset.marketing_campaign_data`
WHERE
  campaign_name = '{{campaign_name}}'
GROUP BY
  date
ORDER BY
  date ASC;
```
```
-- Identify the best and worst performing ads by ROAS or CPA
(SELECT ad_name, SAFE_DIVIDE(SUM(revenue), SUM(spend)) AS roas FROM `your_project.your_dataset.marketing_campaign_data` WHERE campaign_name = '{{campaign_name}}' GROUP BY ad_name ORDER BY roas DESC LIMIT 5)
UNION ALL
(SELECT ad_name, SAFE_DIVIDE(SUM(revenue), SUM(spend)) AS roas FROM `your_project.your_dataset.marketing_campaign_data` WHERE campaign_name = '{{campaign_name}}' GROUP BY ad_name ORDER BY roas ASC LIMIT 5);
```
---
## **Output to User: Follow this format in your response**
---
## Marketing Campaign Review: Q4 Holiday Sale
### 📝 Executive Summary
The Q4 Holiday Sale campaign is performing above target, achieving an **Overall ROAS of 4.2:1** against a goal of 3.5:1, and a **CPA of $28** against a target of $35. **Paid Social (Facebook/Instagram) is the primary driver of this success**, especially with video creative targeted at retargeting audiences. **Paid Search (Google Ads) is currently underperforming**, with high CPCs on non-branded keywords diluting overall profitability. Immediate action is recommended to re-allocate 20% of the remaining Paid Search budget to top-performing Paid Social campaigns.
### 💡 Key Insights & Recommendations
-   **Overall Performance**: The campaign is profitable and on track to exceed its revenue goals by an estimated 15%.
-   **Channel Performance**: Paid Social is delivering a **5.8:1 ROAS**, while Paid Search is lagging at **2.1:1 ROAS**.
    -   **Recommendation**: Decrease spend on underperforming Google Ads ad groups and move that budget to Facebook retargeting campaigns.
-   **Audience Performance**: Retargeting audiences are the most profitable segment with a **CPA of $15**, compared to a **$55 CPA** for Prospecting audiences.
    -   **Recommendation**: While Prospecting is necessary for growth, we should refine our Prospecting audiences on Facebook to focus on lookalikes of high-value customers to improve efficiency.
-   **Creative Performance**: Video ads on Facebook/Instagram Reels are the top performers, generating **2x the ROAS of static image ads**.
    -   **Recommendation**: Pause the bottom 10% of static image ads and create new variations of the top-performing video creative.
---
### 📈 Performance Deep Dive
#### Key Performance Indicators (KPIs) at a Glance
| Metric | Current Result | Target | Status |
| :--- | :--- | :--- | :--- |
| **Return On Ad Spend (ROAS)** | **4.2:1** | 3.5:1 | ✅ On Track |
| **Cost Per Acquisition (CPA)** | **$28** | $35 | ✅ On Track |
| **Total Spend** | $50,000 | $100,000 | ⏳ Pacing |
| **Total Revenue** | $210,000 | $350,000 | ⏳ Pacing |
**Charts (ALL REQUIRED, dark background, light text):**
1.  **ROAS by Channel**
    -   A vertical bar chart comparing the ROAS of Paid Search and Paid Social.
    -   A horizontal dashed line indicates the target ROAS of 3.5:1.
    -   Bars are labeled with their ROAS value (e.g., "5.8:1").
2.  **CPA by Audience Type**
    -   A vertical bar chart comparing the CPA for Prospecting and Retargeting audiences.
    -   A horizontal dashed line indicates the target CPA of $35.
    -   Bars are colored green if below target and red if above.
3.  **Campaign ROAS Trend (Daily)**
    -   A line chart showing the daily ROAS over the campaign's duration.
    -   This helps visualize performance trends, decay, or the impact of optimizations.
---
### 🔍 Creative & Ad-Level Analysis
Analysis of individual ads reveals a clear pattern: engaging, short-form video content is dramatically outperforming static assets. Our top ad, "Holiday-Video-Ad-01," has generated over $30k in revenue on its own. Conversely, several text ads on Google are failing to convert, despite high click costs.
**Top 5 & Bottom 5 Ads by ROAS:**
-   A horizontal bar chart showing two groups of bars.
-   **Top 5 Ads**: Green bars extending to the right, labeled with the ad name and its high ROAS.
-   **Bottom 5 Ads**: Red bars extending to the left, labeled with the ad name and its low (or negative) ROAS. This creates a clear visual contrast between winners and losers.

````

{% endcode %}

## Metric Glossary

You wouldn't trust an analyst that regularly reports different numbers than what your well-maintained dashboards show. It's not because the numbers are necessarily wrong, but since they are inconsistent you have to figure out which one is more correct. A metric glossary can help with that. Here you just list all your most important KPIs/metrics with

* name of metric
* short description (maybe with synonyms)
* query on how to calculate the metric

Dot will use those definitions when giving answers.

**Example**

````
This is our metric glossary. 
Always use the definitions specified here to write queries

## ARR: Annual Recurring Revenue
```sql
SELECT sum(contract_value)
FROM opportunities_table
WHERE status = 'closed'
```

## NRR: Net Revenue Retention
...
````

{% hint style="info" %}
You can nest notes in a hierarchy to stay organized.

You can assign access groups to notes to manage different domains or access patterns.
{% endhint %}

{% hint style="success" %}
**CLI access:** You can also manage notes from the terminal with `dot notes list`, `dot notes create`, etc. This is useful for scripting, bulk imports, and CI/CD pipelines. See [CLI & AI Agent Skill](/developers/cli#notes) for details.
{% endhint %}


# Version Control

Sync your context files with a Git provider for version control, audit trail, and AI/CI interoperability

Dot can sync the files it manages — notes, table documentation, relationships, custom skills, and dashboards — with a Git repository. Pick the provider you already use:

* [**GitHub**](/train-dot/version-control/github) — install the Dot Context Sync GitHub App and pick a repo. Webhooks are auto-created for you.
* [**GitLab**](/train-dot/version-control/gitlab) — paste a Personal Access Token (gitlab.com or self-hosted). Webhook is set up manually with a URL and secret Dot displays for you.

You can connect both providers at the same time — every change in Dot fans out to all configured repos in parallel, and webhook events from any provider pull changes back into Dot.

## Why version-control your context

Your context lives in plain Markdown and YAML files. Keeping it in Git gives you:

* **Audit trail** — every change is a commit. `git blame` shows who changed which note and why.
* **Backup + disaster recovery** — restore an accidentally deleted note with `git revert`, or rebuild a fresh server from the repo.
* **AI/coding-agent interop** — Claude Code, Cursor, and GitHub Copilot can read and contribute to the same context Dot uses.
* **Code review** — important notes and skills go through PR review before reaching production.
* **Cross-environment promotion** — point staging at `dev` and prod at `main` of the same repo, then promote with a PR.
* **Bulk edits in your editor** — power users grep/sed across hundreds of notes locally, push, and Dot picks up the changes.

## What gets synced

```
notes/
├── active/          # Active notes
└── inactive/        # Archived notes

data_sources/
├── active/          # Active table documentation
└── inactive/        # Archived table documentation

relationships.yaml   # Table relationships
apps/                # Dashboards as code (.app sources + .app.lock)
skills/              # Custom skills
```

All files use Markdown (`.md`) with YAML frontmatter, or plain YAML.

{% hint style="info" %}
Only context files are synced. Database connections, user accounts, chat history, schedules, and billing settings stay in Dot.
{% endhint %}

## Version history and undo

You don't need a Git provider to get history and undo. Every change to your context is a commit, and Dot keeps that history for you whether or not you've connected GitHub or GitLab.

Open Version History from the sidebar to see the full list of changes, each with a diff. From there you can roll the whole model back to an earlier point. You can also undo a single thing: open a table on the Model page, look at its history, and revert just that table if something went wrong. Reverting doesn't erase anything. Dot records the undo as a new change, so you can always see what happened and when.

Reverting is an admin action by default. Admins can let modelers do it too with the "Can revert version history to a previous commit" permission, under Settings, then Advanced Settings, then Modeler Permissions.

## Workspaces

Each org and each workspace has its own independent connection — switching to a workspace shows the empty Connect panel even if the parent org is connected. This lets you keep separate workspaces pointing at separate repos (or no repo at all).


# GitHub

Sync your context files with GitHub for version control and automation interoperability

GitHub Sync enables bidirectional synchronization between Dot's context files and a GitHub repository.

## Why GitHub Sync Matters

The real power lies in **interoperability**. By storing context in GitHub using open formats (Markdown with YAML frontmatter), you enable:

* **Coding agents** (Claude Code, Cursor, GitHub Copilot) to read and contribute to the same context Dot uses
* **CI/CD pipelines** to validate or transform documentation automatically
* **Team collaboration** through pull requests and code review

Your data documentation becomes a shared asset that multiple AI assistants can leverage.

## What Gets Synced

```
notes/
├── active/          # Active notes
└── inactive/        # Archived notes

data_sources/
├── active/          # Active table documentation
└── inactive/        # Archived table documentation

apps/                # Dashboards as code
├── _shared/         # Org-wide dashboards (.app sources + .app.lock compiled sidecars)
└── <email>/         # Per-user dashboards

skills/              # Custom skills (skills/<name>/SKILL.md plus supporting files)

relationships.yaml   # Table relationships
```

Context files use Markdown (`.md`) with YAML frontmatter, or plain YAML. Dashboards sync as `.app` sources plus their `.app.lock` compiled sidecars, and custom skills as their `SKILL.md` and supporting text files.

{% hint style="info" %}
Context files, dashboards (apps), and custom skills are synced. Database connections, user settings, and chat history remain in Dot.
{% endhint %}

## Setting Up GitHub Sync

### Prerequisites

* Admin access to your Dot organization
* A GitHub account with permission to install GitHub Apps

### Step 1: Connect with GitHub

1. Go to **Settings** > **Version Control** > **GitHub**
2. Click **Connect with GitHub**

<figure><img src="/files/iUKLCmDYEZTxMVMD8guy" alt="GitHub sync panel in Dot settings"><figcaption><p>Click Connect with GitHub to start the setup</p></figcaption></figure>

3. Install the **Dot Context Sync** GitHub App
4. Choose repository access — either **All repositories** or **Selected repositories**

<figure><img src="/files/9GUNFBcax9bBCURrpBGs" alt="GitHub App installation page"><figcaption><p>Choose repository access during installation</p></figcaption></figure>

{% hint style="info" %}
Both "All repositories" and "Selected repositories" work. If you choose "Selected repositories", you'll need to create your sync repository manually on GitHub first (see Step 2).
{% endhint %}

### Step 2: Configure Your Repository

**Select existing repository** (recommended): Select a repository and branch from the list. This works with both "All repositories" and "Selected repositories" access.

{% hint style="success" %}
**Tip**: If you use "Selected repositories" access, create an empty private repository on GitHub first. When you install or update the GitHub App, make sure the repository is included in the App's access list. It will then appear in Dot's repository dropdown.
{% endhint %}

**Create new repository**: Enter a name, choose a branch, and Dot will create a private repo with all your existing context. This option requires the GitHub App to be installed with "All repositories" access.

{% hint style="info" %}
If your repository already has files (like a README), Dot will preserve them. Dot only manages files in the `notes/`, `data_sources/`, `apps/`, and `skills/` paths, plus `relationships.yaml`.
{% endhint %}

### Step 3: Enable Auto-Sync

Toggle **Auto-sync enabled** to automatically push changes when you update context in Dot's UI.

## How Sync Works

**Push (Dot → GitHub)**: Changes in Dot automatically commit to GitHub when auto-sync is enabled. When you first configure a repository, Dot pushes all existing context files immediately.

**Pull (GitHub → Dot)**: Changes in GitHub (commits, merged PRs) automatically sync to Dot via webhooks.

{% hint style="success" %}
**Collaboration workflow**: Team members or AI coding assistants can submit pull requests. Once merged, changes automatically appear in Dot.
{% endhint %}

**Manual sync**: Force push or pull all files from Settings if needed.

## Limitations

* **One repository per workspace.** Dot syncs your production branch. If you use [environments](/train-dot/environments) and leave environment mirroring on, Dot also pushes a branch for each environment, named `dot/env-<slug>`, so you can review changes as pull requests
* **Text files only** in the managed paths — Markdown/YAML context, `.app` dashboard sources with their `.app.lock` sidecars, and skill files (binary or non-UTF-8 assets are skipped)
* **Last write wins** for simultaneous edits (no conflict resolution UI)

**Not synced**: Database credentials, user accounts, chat history, scheduled reports, organization settings.

## Troubleshooting

| Issue                      | Solution                                                                                                                                 |
| -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| Sync not working           | Check Settings > GitHub for connection status and auto-sync toggle                                                                       |
| Permission errors          | Verify the GitHub App has access to your repo in GitHub Settings > Applications                                                          |
| Files not appearing        | Ensure files are in `notes/`, `data_sources/`, `apps/`, `skills/`, or `relationships.yaml` with correct extensions                       |
| Can't create repository    | This requires "All repositories" access. Create the repo manually on GitHub and select it instead                                        |
| Sync disabled unexpectedly | If you remove a repository from the GitHub App's access, Dot automatically disables sync. Re-add the repo and re-enable sync in Settings |

## Security

Repositories created by Dot are private. Uses secure GitHub App tokens with scoped access. All webhook payloads are cryptographically verified.


# GitLab

Sync your context files with GitLab for version control and AI/CI interoperability

GitLab Sync enables bidirectional synchronization between Dot's context files and a GitLab project. It works with **gitlab.com** and **self-hosted GitLab** (Community or Enterprise Edition).

## Setting Up GitLab Sync

### Prerequisites

* Admin access to your Dot organization
* A GitLab account with the `api` scope on a Personal Access Token

### Step 1: Generate a Personal Access Token

In GitLab, go to **User Settings → Access Tokens** and create a new token with the `api` scope.

{% hint style="warning" %}
Make sure you're on **User Settings**, not **Project Settings**. Both pages look identical and use the same `glpat-` prefix, but a project access token authenticates as a bot user with no personal namespace — it can sync to existing projects but cannot create new ones.
{% endhint %}

### Step 2: Connect with GitLab

1. Go to **Settings** > **Version Control** > **GitLab**
2. Paste your token

<figure><img src="/files/zoZwfWJBkdNBSON9IgRO" alt="GitLab sync connect panel"><figcaption><p>Paste a Personal Access Token to connect</p></figcaption></figure>

If you're on self-hosted GitLab, click **Using self-hosted GitLab?** and enter your instance URL (e.g. `https://gitlab.your-company.com`). Leave it blank for gitlab.com.

<figure><img src="/files/fBvJNiUFSbrmojafof46" alt="GitLab self-hosted URL field"><figcaption><p>Enter your instance host for self-hosted GitLab</p></figcaption></figure>

{% hint style="info" %}
The GitLab URL field is the **instance** URL, not a project URL. If you paste a project URL like `https://gitlab.com/yourname/your-project`, Dot normalizes it to the host (`https://gitlab.com`) automatically.
{% endhint %}

3. Click **Connect with GitLab**. Dot validates the token against GitLab's API and stores it.

### Step 3: Configure Your Project

**Select existing project** (recommended): Pick a project from the list and choose a branch. This is the only option for project- or group-scoped tokens.

**Create new project**: Enter a name. Dot creates a private project under your namespace and pushes existing context files to it. This option requires a Personal Access Token (project bot tokens have no namespace they're allowed to write to).

### Step 4: Set Up the Webhook (optional, for automatic pulls)

Unlike GitHub, GitLab webhooks aren't created automatically. To pull changes into Dot when teammates push to GitLab directly, set up a webhook manually:

<figure><img src="/files/s8eRG5JuIn8PhmmWuC0R" alt="GitLab sync configured state showing webhook URL and secret"><figcaption><p>The configured panel surfaces the webhook URL and secret to paste into GitLab</p></figcaption></figure>

1. In GitLab → your project → **Settings → Webhooks → Add new webhook**
2. Copy the **URL** from Dot's panel into GitLab's URL field
3. Copy the **Secret** from Dot's panel into GitLab's "Secret token" field
4. Enable **Push events** (filter to your sync branch if you like)
5. Save

When someone pushes to the configured branch, GitLab sends a signed event to Dot, and Dot pulls only the changed files.

### Step 5: Enable Auto-Sync

Toggle **Auto-sync enabled** to push changes to GitLab automatically when you edit context in Dot.

## How Sync Works

**Push (Dot → GitLab)**: Changes in Dot become commits in GitLab via the Repository Commits API — one atomic commit per change. The first push after configuration ships your existing context.

**Pull (GitLab → Dot)**: Changes pushed to GitLab trigger your configured webhook; Dot pulls the changed files and invalidates its caches so the AI sees the new content on the next call.

**Manual sync**: Use the **Push to GitLab** / **Pull from GitLab** buttons for one-off triggers.

## Token Types

GitLab has three kinds of tokens that all use the `glpat-` prefix:

| Token                 | Where to generate                | Can connect | Can list/sync existing project | Can create new project |
| --------------------- | -------------------------------- | ----------- | ------------------------------ | ---------------------- |
| Personal Access Token | User Settings → Access Tokens    | ✓           | ✓                              | ✓                      |
| Group Access Token    | Group Settings → Access Tokens   | ✓           | ✓ (within group)               | ✗                      |
| Project Access Token  | Project Settings → Access Tokens | ✓           | ✓ (that project only)          | ✗                      |

If Dot detects a project- or group-scoped token, the **Create new project** option is hidden and a notice explains why. Use a Personal Access Token if you want Dot to create the project for you.

## Limitations

* **One project per workspace**, syncing with a single branch
* **Markdown and YAML only** in the configured directories
* **Last write wins** for simultaneous direct edits (no conflict resolution UI)

**Not synced**: Database credentials, user accounts, chat history, scheduled reports, organization settings.

## Troubleshooting

| Issue                                         | Solution                                                                                                                                                                                                                                      |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| "GitLab rejected the token"                   | Make sure the token has the `api` scope and isn't expired                                                                                                                                                                                     |
| "Can't reach `<url>`"                         | Check the instance URL points at the host (e.g. `https://gitlab.com`), not a project page; verify the host is reachable from Dot's server                                                                                                     |
| "responded 404 for /api/v4/user"              | The URL isn't a GitLab instance — set the host, not a specific project URL                                                                                                                                                                    |
| "This token can't create projects"            | You're using a project- or group-scoped token. Use **Select existing project**, or generate a Personal Access Token at User Settings → Access Tokens                                                                                          |
| Branch is "protected" but I can push manually | Configure should still succeed — Dot no longer preemptively rejects protected branches. If a real push fails later because of branch protection, grant the token user push access in **Project → Settings → Repository → Protected branches** |
| Webhook isn't triggering pulls                | Verify the URL and secret in GitLab match what Dot's panel shows; confirm **Push events** is enabled; check the configured branch matches                                                                                                     |

## Self-hosted GitLab notes

* The instance URL is normalized to scheme + host. Trailing paths or project URLs are stripped automatically.
* Make sure the host is reachable from Dot's server (consider firewall rules and outbound network policy).
* Self-hosted instances often run a different default branch (e.g. `master`). Dot uses whatever branch GitLab created the project with — no need to override.

## Security

* The token is stored encrypted at rest alongside other org secrets.
* The webhook secret is generated per-org and rotated when you disconnect/reconnect.
* The webhook handler verifies `X-Gitlab-Token` constant-time and authenticates per-org — a webhook signed with org A's secret can't trigger pull jobs for org B sharing the same `project_id` (e.g. across self-hosted instances).
* Non-admin status callers can see the connection state but never the webhook secret.


# Environments

Develop and test changes to Dot's knowledge in an isolated environment — then merge them to production when they're ready

Environments let you change Dot's knowledge — table documentation, notes, relationships, skills — without touching what your colleagues see. Each environment is an isolated copy of your data model, backed by its own git branch. You switch in, make changes, test them against real questions, and merge back to production in one click.

If you work with dbt, this will feel familiar: an environment is to Dot what a dev target is to dbt. You can even point an environment at your dbt dev schema, so Dot and dbt develop against the same data.

<figure><img src="/files/mAdZHTuo4sPgUqnsSpdx" alt="Environment switcher in the sidebar"><figcaption><p>Switch environments from the bottom of the sidebar — Production stays untouched while you work</p></figcaption></figure>

**Why this matters**: Your team relies on Dot's answers. Editing documentation live means every half-finished description and experimental relationship immediately shapes production answers. Environments give you a place to get it right first:

* **Safe iteration** — remodel a domain, rewrite descriptions, or test new notes while production answers stay stable.
* **Test with real questions** — chat with Dot inside the environment and verify answers before anyone else is affected.
* **dbt-style workflows** — point the environment at your dev schema/database, develop dbt models and Dot docs together, and promote both when ready.
* **Reviewable changes** — every environment is a git branch. See the exact file diff before merging, just like a pull request.

{% hint style="info" %}
Environments are available to **admins and modelers**. Regular users always see production.
{% endhint %}

## What's isolated, what's shared

| Isolated per environment                                    | Shared with production                     |
| ----------------------------------------------------------- | ------------------------------------------ |
| Table documentation, notes, relationships, reports, skills  | Database connections (overridable)         |
| Warehouse target overrides (dev schema/database)            | Users, permissions, and groups             |
| [Root](/train-dot/context-agent) sessions and their changes | Chat history (tagged with the environment) |

Chats you run inside an environment appear in the regular History page, tagged with the environment's name — so usage stays visible in one place.

## Create and switch

1. Click the environment switcher at the bottom of the sidebar (it shows **Production** by default).
2. Choose **Manage environments**.
3. Name the environment, pick a color, and click **Create environment**. The new environment forks from production (or from another environment via **Fork from**).
4. Click **Switch** on the environment to start working in it.

<figure><img src="/files/UNxRbCEhHI1BI9RV5mCQ" alt="Environment manager"><figcaption><p>Create, switch, diff, merge, and delete environments in one place</p></figcaption></figure>

While an environment is active, every page shows a slim banner with the environment's color so you always know where you are. The switcher in the sidebar shows the same colored dot.

<figure><img src="/files/iM45tXJvqVGo4qKws6cd" alt="Environment banner"><figcaption><p>The banner at the top tells you this tab works in <code>env/dev_rick</code>, isolated from production</p></figcaption></figure>

Environments are per-tab: switching in one browser tab doesn't affect your other tabs, so you can keep production open side by side.

## Quick fixes from a chat

You don't need to create an environment for small fixes. When you start a [Root](/train-dot/context-agent) session from production, Dot automatically creates a throwaway environment behind the scenes. Root makes its changes there, you verify the result in the chat, and when you apply the changes to production the throwaway environment is cleaned up automatically.

This means production documentation is never edited in place — every change, however small, goes through an isolated branch.

A throwaway environment is disposable on purpose. Dot deletes it after seven days without activity, and it never gets a branch in your Git repository, even when you have environment mirroring on.

To hold on to one, open **Manage environments** and click **Keep it** on the environment. Dot then treats it like any environment you created yourself: it stays until you delete it, Root leaves it alone when the chat finishes, and it gets its own `dot/env-<slug>` branch if environment mirroring is on.

## Work against your dbt dev target

By default an environment reads from the same warehouse schemas as production. For dbt-style development you can point it at your **dev target** instead:

1. In **Manage environments**, click **Warehouse targets** on the environment.
2. Enter the same values your dbt dev target uses (for example schema `dbt_demo_dev` instead of `dbt_demo` — for Snowflake: database/schema, for BigQuery: project/dataset).
3. Click **Save targets**.

<figure><img src="/files/QdZ7QhzZmNtjGI2WUHjN" alt="Warehouse target editor"><figcaption><p>Point the environment at your dbt dev schema — queries run against it while production stays untouched</p></figcaption></figure>

From now on, questions asked inside this environment run against your dev schema. Dot follows dbt's defer semantics: **tables that exist in your dev schema are read from there; everything else falls back to production**. You can run `dbt build` on just the model you're changing — exactly like `dbt build --defer`.

After you've built or changed models in the dev schema, click **Sync from dev target** (or run `dot env sync-target`) to refresh the environment's table documentation from the dev schema — new columns and new tables show up in the environment only.

{% hint style="warning" %}
Target overrides redirect **queries and metadata sync** for that environment only. The connection itself — credentials, host, warehouse — is still shared with production.
{% endhint %}

## Review and merge

When the work is ready:

1. Open **Manage environments** and click **Diff** to see every file the environment changed compared to production. Each file is marked `added`, `modified` or `deleted`.
2. Click a file to read the change itself, line by line, with additions in green and removals in red. One file stays open at a time, so click another to switch. You don't have to leave Dot to see what a merge would do.
3. Click **Merge to production** and confirm. Dot merges the environment's branch into the production model.
4. Optionally delete the environment after merging — or keep it for the next iteration.

<figure><img src="/files/AfXsKziJ25ZVa93WXFQR" alt="Environment diff"><figcaption><p>The diff lists every documentation file the environment changed</p></figcaption></figure>

<figure><img src="/files/RtRQZfIPEfAGm6XJA1ZE" alt="An expanded file in the environment diff"><figcaption><p>Click a file to read its changes line by line</p></figcaption></figure>

Admins can always merge. Modelers need the "Can merge changes to production" permission, which admins grant under Settings, then Advanced Settings, then Modeler Permissions.

### Promote with a pull request instead

If you connect a Git provider and want changes to go through review, you don't have to merge inside Dot. You can open a pull request instead. On the environment, click **Open pull request**, and Dot opens a PR from the environment's branch into your production branch on GitHub or GitLab. Your team reviews the diff there and merges when it's ready, and Dot picks up the change. Anyone can open a pull request. Merging to production is the part that needs the permission.

{% hint style="info" %}
This works when Dot is mirroring environments to Git, which it does by default. Each lasting environment gets its own branch in the same repository as production, named `dot/env-<slug>`. You can turn this off with the "Mirror environments to Git" toggle under Settings, then Version Control. The throwaway environments Root creates for quick fixes aren't mirrored, unless you keep one.
{% endhint %}

## For coding agents: CLI & API

Environments are fully scriptable, which makes them the natural unit of work for AI coding agents: create an environment, make changes, verify, merge — without ever touching production. The [Dot CLI](/developers/cli) ships an `env` command group:

```bash
dot env list                          # Production + all environments
dot env create "dev-rick" --color "#3A86E8"
dot env use dev-rick                  # all following CLI commands run in this env
dot env target set dev-rick db.getdot.ai:5432:db --schema dbt_demo_dev
dot env sync-target dev-rick          # refresh env docs from the dev schema
dot env diff                          # changed files vs production
dot env conflicts                     # predict merge conflicts
dot env merge dev-rick --confirm      # promote to production
dot env delete dev-rick --confirm
```

The active environment (`dot env use`) is sent as an `X-Dot-Environment` header on every request. Override it per command with `--env <id>` or the `DOT_ENV` environment variable.

Using the [REST API](/developers/api) directly, send the same header with the environment's id:

```bash
# Everything below is scoped to the environment — production is untouched
curl -H "X-API-KEY: $DOT_API_KEY" -H "X-Dot-Environment: $ENV_ID" \
  -H "Content-Type: application/json" https://app.getdot.ai/api/agentic \
  -d "{\"messages\": [{\"role\": \"user\", \"content\": \"Which columns does dim_customers have?\"}], \"chat_id\": \"$(uuidgen)\"}"
```

| Endpoint                                     | Purpose                                       |
| -------------------------------------------- | --------------------------------------------- |
| `GET /api/environments`                      | List environments                             |
| `POST /api/environments`                     | Create (`{"name", "color", "source_env_id"}`) |
| `GET /api/environments/{id}/diff`            | Changed files vs production                   |
| `GET /api/environments/{id}/conflicts`       | Predict merge conflicts                       |
| `PUT /api/environments/{id}/targets`         | Set warehouse target overrides                |
| `POST /api/environments/{id}/sync_target`    | Sync env docs from the dev target             |
| `POST /api/environments/{id}/merge`          | Merge to production (`{"confirm": true}`)     |
| `DELETE /api/environments/{id}?confirm=true` | Delete                                        |

A typical agent recipe — fix documentation for a model you just changed in dbt:

```bash
dot env create "fix-customer-tier" && dot env use fix-customer-tier
dot env target set fix-customer-tier <connection-id> --schema dbt_dev_alice
dbt build --select dim_customers     # build the model into the dev schema
dot env sync-target fix-customer-tier
dot ask "Which columns does dim_customers have?"   # verify Dot sees the dev version
dot env merge fix-customer-tier --confirm --delete-after
```


# Permissions

permissive or restrictive, however you like it

## Roles in Dot

There are three roles that control what you can do in Dot:

* **Admin** — Full access to settings, connections, users, billing, and model management.
* **Modeler** — Can manage the data model, run evaluations, and view chat history for their groups. No access to settings, billing, or user management. Admins can adjust what a modeler can do (see [Modeler permissions](#modeler-permissions) below).
* **User** — Can chat and ask questions only.

Roles are managed by admins on the Settings > Users page. Hover over the info icon next to the Role column to see a summary of each role.

<figure><img src="/files/mrZ1P3fUZiQNoSBt7mTj" alt=""><figcaption><p>Users page with role descriptions tooltip</p></figcaption></figure>

### Access Matrix

| Feature         | Admin | Modeler | User |
| --------------- | :---: | :-----: | :--: |
| Chat            |  Yes  |   Yes   |  Yes |
| History         |  Yes  |   Yes   |   —  |
| Model page      |  Yes  |   Yes   |   —  |
| Evaluation      |  Yes  |   Yes   |   —  |
| Settings        |  Yes  |    —    |   —  |
| Connections     |  Yes  |    —    |   —  |
| Billing         |  Yes  |    —    |   —  |
| User management |  Yes  |    —    |   —  |

The Modeler row shows the defaults. That row isn't fixed, though. Admins can change what a modeler can do (see the next section). One thing worth calling out: a modeler's History only shows chats from people in groups they share, not everyone's chats, and admins can hide the History page from modelers entirely.

## Modeler Permissions

The Modeler role is a starting point, not a fixed set of rules. Admins shape it under Settings, then Advanced Settings, then Modeler Permissions.

A few things are on by default, because modelers have always been able to do them. You can turn any of them off:

* See and edit the Skills tab
* See the History page (limited to their own groups)
* Chat without restrictions

Everything else is off by default. An admin can grant it when a modeler needs it:

* Trigger and schedule connection syncs
* Assign workspace users to groups
* Merge changes to production
* Review Root's proposals, meaning see them and merge or reject them
* Revert version history to an earlier commit
* Manage all apps, like an admin

This way the role stays tight by default, and you hand out exactly the extra access a person needs. These settings are per workspace, so a modeler can have different access in different workspaces.

## Data Access Control

Groups manage who gets access to what data. A user can only query a table if they share at least one group with that table.

* A user can belong to multiple groups and a table can also belong to multiple groups.
* By default, new tables are assigned the group `all_users` and new users are also added to this group.
* An admin can configure that new users don't automatically get the `all_users` group if your organization operates on a "need to know" basis.

To manage groups for a table, open the table on the Model page and click the **Access** tab.

<figure><img src="/files/eaI6XxQCxS9YegbMH2pJ" alt=""><figcaption><p>Access groups for a table on the Model page</p></figcaption></figure>

## Row Level Permissions

Groups can also be used to apply row-level filtering on tables.

To enable row level permissions, open a table on the Model page, go to the **Access** tab, enable the **Row Level Permissions** toggle, and specify a where clause template.

The where clause should contain the placeholder `${groupname}`, which is replaced by the user's first group name at query time. You can also use `${all_groupnames}` to match against all of a user's groups.

**Examples:**

```sql
-- Filter by a single group
WHERE '${groupname}' = country_code

-- Map group names to values
WHERE CASE '${groupname}' WHEN 'germany' THEN 'DE' WHEN 'france' THEN 'FR' END = country_code

-- Match any of the user's groups
WHERE country_code IN (${all_groupnames})
```

<figure><img src="/files/zYRWPEAufOdQ5JAGvzQw" alt=""><figcaption><p>Row Level Permissions configuration on the Access tab</p></figcaption></figure>


# User Feedback

is valuable input for learning

Users can either upvote or downvote a response from Dot.

This feedback is used to expand Dots knowledge base, and give admins relevant insights to manage Dot.

## Positive Feedback

"Strengthen strengths" made young Boris Becker a world champion in Tennis 🎾.\
When a user clicks 👍, Dot will store the generated query for this question and reuse it for similar questions in the future.

<figure><img src="/files/DScrxLOpD2gyi0C1Klst" alt=""><figcaption><p>Click thumbs up to save a successful query</p></figcaption></figure>

{% hint style="info" %}
Before a generated query gets used by Dot an admin needs to select it on the Model page.
{% endhint %}

## Negative feedback

👎 is a signal for admins that Dot's knowledge base needs to get adjusted to be better able to answer this question.

## Admin Overview

Admins can select to see the history of all users. All conversations with at least 1 dislike will show up with `problem` and if there is no dislikes and at least 1 like it shows as `success`.

<figure><img src="/files/ceaSGUFIsmqE1nfEjdZ6" alt=""><figcaption><p>View feedback across all users in History</p></figcaption></figure>


# Workspaces

Separate environments for different teams

Workspaces are isolated environments within your organization—separate users, data connections, and permissions, but shared billing.

**Use cases**: Team separation (Sales vs Finance), client isolation, regional data separation, testing environments.

### Creating a Workspace

*Org admins only*

1. **Settings → Workspaces → Create Workspace**
2. Enter a name, optionally copy data from another workspace
3. Done

<figure><img src="/files/1dcnOdmOBqES7kwy9OLo" alt=""><figcaption><p>Create a new workspace with optional data copying</p></figcaption></figure>

Limits: 10 workspaces (free) / 200 (unlimited).

### Switching Workspaces

Click your **workspace name** (bottom-left) → select another workspace.

<figure><img src="/files/qCxUxDtzBzLkzWxIHfwK" alt=""><figcaption><p>Switch between workspaces from the sidebar</p></figcaption></figure>

You can set a default workspace in your user settings.

### Adding Users

1. **Settings → Workspaces** → find workspace → **Manage Users**
2. Enter email, select role (User/Admin), click Add

New users invited to a workspace can only see that workspace. Admins can grant full org access later.

### Slack & Teams Routing

Route specific channels to workspaces. See [Channel Routing](/integrations/slack-and-teams/channel-routing).


# Custom Skills

Extend what Dot can do.

> This is a premium feature. Please contact us for access.

Custom skills allows you to teach new things to Dot. It's mostly intended for connecting to external systems. Here are some of the things you might want to do.

<details>

<summary>Query Another Dot Instance</summary>

This example shows how to query data from another Dot instance using its API. The response will be properly formatted for Dot's custom skill parser.

Paste this code into custom skills in Model > Custom Skills > Add skill

The parameters are:

* user\_request (String): The user's question
* chat\_id (String): The conversation ID (use "new" for new conversations)

**Description**

Please add this description to the skill along with whatever else you want to add.

```
The chat_id parameter determines the conversation state:
- If chat_id is "new", it will start a new chat
- If you provide an existing chat_id, it will continue that conversation

When a DataFrame is returned:
- It's automatically stored in your  memory
- You can use it like a normal DataFrame to chain with other tools
- For example: visualize it or display it

Example workflow:
1. First call: "Show me sales by region" (chat_id="new")
   → Returns data and a chat_id
2. Follow-up: "Now show only regions over $1M" (use returned chat_id)
   → Continues the same conversation context

```

**Code**

```python
import requests
import uuid
import os
from dotenv import load_dotenv

load_dotenv()

API_KEY = os.getenv("DOT_API_KEY")  # Get from Settings > API Tokens in your Dot instance
BASE_URL = "https://app.getdot.ai/api"  # Or "https://eu.getdot.ai/api" for EU

headers = {"API-KEY": API_KEY, "Content-Type": "application/json"}

# These variables are injected by Dot when running as a custom skill:
# - user_request: The user's question
# - chat_id: Conversation ID (use "new" or None for new conversations)

# For local testing, uncomment these:
user_request = "Show me total sales by product category"
chat_id = "new"


try:
    if "chat_id" not in locals() and "chat_id" not in globals():
        chat_id = "new"

    is_new_chat = chat_id is None or chat_id == "" or chat_id == "new"

    if is_new_chat:
        chat_id = str(uuid.uuid4())
        url = f"{BASE_URL}/agentic"
        payload = {"messages": [{"role": "user", "content": user_request}], "chat_id": chat_id}
    else:
        url = f"{BASE_URL}/agentic_with_history"
        payload = {"new_message": {"role": "user", "content": user_request}, "chat_id": chat_id}

    # The call returns the whole conversation once Dot has finished answering
    response = requests.post(url, headers=headers, json=payload, timeout=600)
    response.raise_for_status()
    messages = response.json()

    # The answer is the last message with a formatted_result — a list of parts,
    # of which the text ones joined together are what a person reads
    for message in reversed(messages):
        parts = (message.get("additional_data") or {}).get("formatted_result") or []
        texts = [str(p["data"]) for p in parts if p.get("type") == "text" and p.get("data")]
        if texts:
            print("\n\n".join(texts))
            break

    print(f"\nTo continue this conversation, use chat_id: {chat_id}")

except requests.exceptions.ConnectionError:
    print("Error: Could not connect to Dot API. Please check your BASE_URL and network connection.")
except requests.exceptions.HTTPError as e:
    if e.response.status_code == 401:
        print("Error: Invalid API key. Please check your API_KEY in Settings > API Tokens.")
        print(e)
    elif e.response.status_code == 404:
        print("Error: API endpoint not found. Please check your BASE_URL.")
        print(e)
    else:
        print(f"HTTP Error {e.response.status_code}: {e.response.text}")
except requests.exceptions.RequestException as e:
    print(f"Error calling Dot API: {str(e)}")
except Exception as e:
    print(f"Unexpected error: {str(e)}")
 
```

</details>

<details>

<summary>Create a Notion ticket</summary>

Opens a Notion ticket whenever a specific condition is met and specified by the customer

### **What gets created in Notion**

**Name** → short summary of the task (Dot always summarizes)

**Description** → brief context + next steps (Dot always writes both)

**Status** → one of: Not started | To do | In progress | Done, or the ones you need

**Due Date** → format YYYY-MM-DD

**Priority** → must match an existing option in your Notion DB

### **What you see in chat**

```
Notion ticket created succesfully ✅ - You can review the notion ticket <a href="{{page_url}}">here</a> and all tickets in <a href="{{database_url}}">here</a>
```

The first link opens the new page, the second opens the database. You can also personalize this in the Dot description.

### **Inputs Dot needs**

**name** → short summary (action-oriented)

**description** → context + next steps

**status\_name** → Not started | To do | In progress | Done

**priority** → existing Priority option in your DB

**due\_date** → YYYY-MM-DD

### **How it works behind the scenes**

Skill maps everything into a single main() that uses the Notion API to create the page and fetch the page/database URLs

Your Notion token and database ID are configured in the Dot skill setup window (no local setup needed)

### **Example**

**name:** Reach out to ACME about low weekly queries

**description:** Usage dipped below threshold. Next steps: email champion with best-practice guide; schedule 20-min optimization session

**status\_name:** To do

**priority:** High

**due\_date:** 2025-08-25

Dot creates the ticket with those fields and replies with the two links

### Dot Skill Description

```markdown
Open a new notion ticket if the usage of Dot for a customer is too low. Everything should be mapped to the main() function.

You should always summarize the task as part of the name parametr, and create a brief description and next steps for the description parameter.

You should consider for the status parameter the following available options: "Not started", "To do", "In progress", "Done".

You should always consider dates to have the format "YYYY-MM-DD".

For the message output to the user, please retrieve both URLs for the Notion Page and Database for easy access.

The message output should look like this:

"Notion ticket created succesfully ✅ - You can review the notion ticket here and all tickets in here" - You should include the links to the page and database as part of the a tags in the first and second "here" words that are part of the message. They should be clickable for better experience.
```

### Python Function

```python
#!/usr/bin/env python3
"""
Single-file Notion API integration script.

This script provides a complete Notion API wrapper with helper functions
for creating pages in Notion databases. Uses only the requests library.

Usage:
    Set environment variables:
    - NOTION_API_TOKEN: Your Notion integration token
    - NOTION_DATABASE_ID: Target database ID (optional, can be passed directly)
    
    Then run: python notion_integration.py
    
    Or import and use the functions in your automation:
    from notion_integration import NotionWrapper, title_prop, rich_text
"""

import json
import logging
import os
import sys
from datetime import datetime
from typing import Any, Dict, List, Optional, Union

import requests

# Configure logging
def setup_logging():
    """Set up logging configuration."""
    level = os.getenv("LOG_LEVEL", "INFO").upper()
    logging.basicConfig(
        level=getattr(logging, level, logging.INFO),
        format="%(asctime)s | %(levelname)-8s | %(message)s",
    )

class NotionWrapperError(Exception):
    """Custom exception for Notion wrapper operations."""
    pass

class NotionWrapper:
    """Simple Notion API wrapper using only requests."""
    
    def __init__(self, token: Optional[str] = None):
        """Initialize with Notion API token."""
        self.token = token or os.getenv('NOTION_API_TOKEN')
        if not self.token:
            raise NotionWrapperError("NOTION_API_TOKEN required")
        
        self.base_url = "<https://api.notion.com/v1>"
        self.headers = {
            "Authorization": f"Bearer {self.token}",
            "Notion-Version": "2022-06-28",
            "Content-Type": "application/json"
        }
        logging.info("NotionWrapper initialized")
    
    def create_page(self, database_id: str, properties: Dict[str, Any]) -> str:
        """Create a new page in a Notion database."""
        if not database_id or not properties:
            raise NotionWrapperError("database_id and properties required")
        
        url = f"{self.base_url}/pages"
        data = {
            "parent": {"database_id": database_id},
            "properties": properties
        }
        
        try:
            response = requests.post(url, headers=self.headers, json=data)
            
            if response.status_code != 200:
                # Get detailed error information
                try:
                    error_details = response.json()
                    error_msg = f"Notion API Error {response.status_code}:\\n"
                    error_msg += f"Message: {error_details.get('message', 'No message provided')}\\n"
                    if 'code' in error_details:
                        error_msg += f"Code: {error_details['code']}\\n"
                    if 'details' in error_details:
                        error_msg += f"Details: {error_details['details']}\\n"
                    error_msg += f"\\nRequest data sent:\\n{json.dumps(data, indent=2)}"
                except:
                    error_msg = f"HTTP {response.status_code}: {response.text}\\n"
                    error_msg += f"Request data sent:\\n{json.dumps(data, indent=2)}"
                
                logging.error(error_msg)
                raise NotionWrapperError(error_msg)
            
            page_id = response.json()["id"]
            logging.info(f"Created page: {page_id}")
            return page_id
            
        except requests.exceptions.RequestException as e:
            error_msg = f"Request failed: {e}"
            logging.error(error_msg)
            raise NotionWrapperError(error_msg)

# Helper functions for property types
def title_prop(text: str) -> Dict[str, Any]:
    """Create title property."""
    return {"title": [{"text": {"content": text}}]}

def rich_text(text: str) -> Dict[str, Any]:
    """Create rich text property."""
    return {"rich_text": [{"text": {"content": text}}]}

def select(name: str) -> Dict[str, Any]:
    """Create select property."""
    return {"select": {"name": name}}

def multi_select(names: List[str]) -> Dict[str, Any]:
    """Create multi-select property."""
    return {"multi_select": [{"name": name} for name in names]}

def status(name: str) -> Dict[str, Any]:
    """Create status property."""
    return {"status": {"name": name}}

def date(start: str, end: Optional[str] = None) -> Dict[str, Any]:
    """Create date property."""
    return {"date": {"start": start, "end": end}}

def number(value: Union[int, float]) -> Dict[str, Any]:
    """Create number property."""
    return {"number": value}

def checkbox(checked: bool) -> Dict[str, Any]:
    """Create checkbox property."""
    return {"checkbox": checked}

def url(link: str) -> Dict[str, Any]:
    """Create URL property."""
    return {"url": link}

def email(address: str) -> Dict[str, Any]:
    """Create email property."""
    return {"email": address}

def phone_number(number: str) -> Dict[str, Any]:
    """Create phone number property."""
    return {"phone_number": number}

def relation(page_ids: List[str]) -> Dict[str, Any]:
    """Create relation property."""
    return {"relation": [{"id": page_id} for page_id in page_ids]}

def get_page_url(page_id: str) -> str:
    """Get the actual page URL from Notion API."""
    try:
        token = os.getenv('NOTION_API_TOKEN')
        headers = {
            "Authorization": f"Bearer {token}",
            "Notion-Version": "2022-06-28",
        }
        response = requests.get(f"<https://api.notion.com/v1/pages/{page_id}>", headers=headers)
        response.raise_for_status()
        page_data = response.json()
        return page_data.get('url', '')
    except requests.exceptions.RequestException as e:
        logging.error(f"Could not retrieve page URL: {e}")
        return ""

def get_database_url(database_id: str) -> str:
    """Get the actual database URL from Notion API."""
    try:
        token = os.getenv('NOTION_API_TOKEN')
        headers = {
            "Authorization": f"Bearer {token}",
            "Notion-Version": "2022-06-28",
        }
        response = requests.get(f"<https://api.notion.com/v1/databases/{database_id}>", headers=headers)
        response.raise_for_status()
        db_data = response.json()
        return db_data.get('url', '')
    except requests.exceptions.RequestException as e:
        logging.error(f"Could not retrieve database URL: {e}")
        return ""

def main(name, description, status_name, priority, due_date):
    """Main function for testing."""
    setup_logging()
    
    # Check for required environment variables
    db_id = os.getenv("NOTION_DATABASE_ID")
    if not db_id:
        print("❌ NOTION_DATABASE_ID not set")
        print("Set it with: export NOTION_DATABASE_ID='your_database_id'")
        sys.exit(1)
    
    token = os.getenv("NOTION_API_TOKEN")
    if not token:
        print("❌ NOTION_API_TOKEN not set")
        print("Set it with: export NOTION_API_TOKEN='your_token'")
        sys.exit(1)
    
    try:
        # Create page with only Name property (since that's what your database has)
        notion = NotionWrapper()
        properties = {
            "Name": title_prop(f"{name}"),
            "Description": rich_text(description),
            "Status": status(status_name),
            "Due Date": date(due_date),
            "Priority": select(priority)
        }
        page_id = notion.create_page(db_id, properties)
        
        # Generate URLs
        page_url = get_page_url(page_id)
        database_url = get_database_url(db_id)
        
        print("\\n🎉 Success! Created Notion page:")
        print(f"📝 Page ID: {page_id}")
        print(f"🔗 Page URL: {page_url}")
        print(f"📊 Database URL: {database_url}")
        print("\\nYou can click these URLs to view in Notion!")
        
        return {
            "page_id": page_id,
            "page_url": page_url,
            "database_id": db_id,
            "database_url": database_url
        }
        
    except NotionWrapperError as e:
        print(f"❌ Error: {e}")
        sys.exit(1)

if __name__ == "__main__":
    main(name, description, status_name, priority, due_date)

```

</details>

### Common Use Cases

1. **Workflow automation** based on data analysis results from Dot: Create Jira/Linear tickets, update Notion docs, send Slack/email messages with findings from Dot.
2. **Enrich Analysis with Additional Data:** Connect data that is not available in your data warehouse, like customer info from a CRM or real-time data from a WMS.
3. **Advanced Analysis Beyond SQL:** Run small ML models inside the sandbox, or host them on a server behind an API that Dot can call. This opens up a whole new world of ML and other analysis. For example forecasting important metrics on the fly.

### **How Custom Skills Work**

Custom skills appear as natural extensions of Dot's capabilities. Users simply ask questions in plain language, and Dot automatically:

1. Identifies when a custom skill is relevant based on the query
2. Extracts required parameters from the conversation context
3. Executes the skill and presents results accordingly.

### **Creating Custom Skills**

Skills live on the Skills page. Open the Model page and go to the Skills tab, or go straight to `/skills`.

A skill is mostly a set of written instructions. You describe, in plain words, what the skill does and when Dot should use it, and Dot follows those instructions the same way it follows your notes. If the skill needs to run code, you add a Python script next to the instructions.

To add one, click **Add skill** and fill in:

* A short name, for example `add_issue_to_jira`.
* A description of what it does and when to use it. Be specific. This is what helps Dot pick the right skill for a question.
* The instructions, written as plain markdown. Explain the steps, mention any script Dot should run, and call out common mistakes so it avoids them.
* Any Python scripts the skill needs.
* Secrets, like API keys, stored as named environment variables. Dot passes them in when the skill runs, so nothing sensitive sits in the instructions.

Two toggles control how a skill behaves. **Active** turns it on or off. **Network** decides whether the script can reach the internet, and it's off by default, so a skill can't make outside calls unless you allow it. A skill that talks to a service like Jira or Notion needs it on. You can also limit a skill to certain user groups if not everyone should use it.

### Install a ready-made skill

You don't have to write everything yourself. The Skills page has a small marketplace of prebuilt skills you can add with one click, including skills for working with PDF, Word, Excel, and PowerPoint files. You can also install a skill from a link someone shares with you. Under the hood every skill is just a `SKILL.md` file plus any supporting scripts, so you can download one, adjust it, and upload it again.

### **Technical Architecture**

#### **Execution Environment**

* **Isolation**: Each execution runs in a Docker container with process-level isolation and resource limitations.
* **Timeout**: 600 seconds maximum per execution
* **Python Version**: 3.12

#### **Available Packages**

Common data packages come pre-installed, including `pandas`, `numpy`, `requests`, `scikit-learn`, `xgboost`, `prophet`, `pyarrow`, `python-pptx`, and `python-dotenv`. If you need something else, get in touch and we'll help.

### Best Practices

#### 1. Parameter Validation

There is a chance Dot makes a mistake. Try to catch as many errors as possible programmatically. Be defensive and provide good feedback. Validation need not be limited to just types—it can be more complex (e.g., len(df) < 1000, check if combinations of parameters are valid).

```python

# process_data function
# argument: df: pd.DataFrame, threshold: float

  # Validate dataframe
  if df.empty:
      print("Error: Empty dataframe provided")

  required_columns = ['amount', 'date', 'category']
  missing = [col for col in required_columns if col not in df.columns]
  if missing:
      print(f"Error: Missing columns: {missing}")

  # Validate threshold
  if not 0 <= threshold <= 1:
      print("Error: Threshold must be between 0 and 1")

```

#### 2. Result Handling

You can pass back results using the `print()` statement.

You can pass back a dataframe to Dot by doing:

`print(dataframe)` — we will automatically handle the logic to convert this into a format that Dot can understand.

Currently we only support print statements with one argument inside (print(a,b) will not work).

Structure your print statements to be clear and easily understandable for Dot. Include only the relevant info and no unnecessary logs.

If something goes wrong during execution, handle it gracefully. Be explicit. Always try to pass back the status of the tool execution (complete/partial/failed). If failed or partial, provide feedback to Dot on what went wrong and how to fix it.

<details>

<summary>Example Python script with explicit error handling</summary>

```python

# fetch_customer_data_from_crm(account_name: str)

api_key = os.environ.get('CRM_API_KEY')
if not api_key:
    print("Status: FAILED - CRM_API_KEY not configured in skill secrets")
    exit()

# Validate input
if 'account_name' not in locals():
    print("Status: FAILED - Account name is required")
    exit()

if not account_name:
    print("Status: FAILED - Account name is required")
    exit()

if not isinstance(account_name, str):
    print("Status: FAILED - Account name must be a string")
    exit()

try:
    # Search for account
    search_url = f"<https://api.crm.com/accounts?name={account_name}>"
    search_response = requests.get(
        search_url,
        headers={"Authorization": f"Bearer {api_key}"}
    )

    if search_response.status_code != 200:
        print(f"Status: FAILED - CRM API error {search_response.status_code}")
        exit()

    accounts = search_response.json()

    if not accounts:
        print(f"Status: PARTIAL - No account found matching '{account_name}'. Try a different spelling or partial name.")
        exit()

    if len(accounts) > 1:
        # Multiple matches - return what we found
        account_list = [f"- {acc.get('name', 'N/A')} (ID: {acc.get('id', 'N/A')})" for acc in accounts[:5]]
        output = (
            f"Status: PARTIAL\\\\n"
            f"Found {len(accounts)} accounts matchIf something goes wrong during execution, handle it gracefully. Be explicit. Always try to pass back the status of the tool execution (complete/partial/failed). If failed or partial, provide feedback to Dot on what went wrong and how to fix it.ing '{account_name}':\\\\n"
            f"{chr(10).join(account_list)}\\\\n"
            f"Please be more specific in your query."
        )
        print(output)
        exit()

    # Single match - get full details
    account_id = accounts[0].get('id')
    if not account_id:
        print("Status: FAILED - Account data missing 'id' field.")
        exit()

    details_url = f"<https://api.crm.com/accounts/{account_id}/details>"
    details_response = requests.get(
        details_url,
        headers={"Authorization": f"Bearer {api_key}"}
    )

    if details_response.status_code == 200:
        account_data = details_response.json()

        # Convert to dataframe for analysis
        df = pd.DataFrame([account_data])
        print(df)  # Make dataframe available to agent

        summary = (
            f"Status: COMPLETE\\\\n"
            f"Retrieved data for {account_data.get('name', 'N/A')}\\\\n"
            f"Industry: {account_data.get('industry', 'N/A')}\\\\n"
            f"Annual Revenue: ${account_data.get('annual_revenue', 0):,.2f}\\\\n"
            f"Employee Count: {account_data.get('employees', 'N/A')}\\\\n"
            f"Full details available in dataframe"
        )
        print(summary)

    else:
        output = (
            f"Status: PARTIAL\\\\n"
            f"Found account {accounts[0].get('name', 'N/A')} but could not fetch details\\\\n"
            f"Error: API returned {details_response.status_code}\\\\n"
            f"Basic info: ID={account_id}"
        )
        print(output)

except requests.exceptions.Timeout:
    print("Status: FAILED - Request timed out. API may be slow or unavailable.")
except requests.exceptions.ConnectionError:
    print("Status: FAILED - Could not connect to CRM API. Check network settings.")
except ValueError as e:
    print(f"Status: FAILED - Invalid JSON response from API: {str(e)}")
except Exception as e:
    print(f"Status: FAILED - Unexpected error: {str(e)}")

```

</details>


# Agent Compute Credits

How Dot measures usage — and where to find current pricing

Dot usage is measured in **Agent Compute Credits (ACCs)**. An ACC reflects the amount of work Dot does to answer a request: the queries it runs, the reasoning steps it takes, and the visualizations it builds. A quick lookup uses very little; a deep, multi-step [investigation](/using-dot/analyze) uses more — and a higher [energy mode](/using-dot/analyze#energy-modes) does more work per question.

Because Dot and the underlying models improve continuously, the same question can take a different path and a different amount of work over time. Treat any figure as a guide, not a guarantee.

{% hint style="info" %}
For current per-plan credit allowances and rates, see the [Dot pricing page](https://www.getdot.ai/pricing).
{% endhint %}

### Keeping usage predictable

Admins can cap how many ACCs a single user or a group can consume, so spend stays predictable and no one user draws down the pool unexpectedly. You configure these limits in Dot's settings.

There's also a built-in guardrail for single answers. Every now and then a hard question makes Dot do a lot of work at once. So this doesn't quietly run up a big bill, Dot pauses partway through and asks whether to keep going once an answer passes a set amount of credits, about 30 by default. You'll see this in the app, and in Slack, Teams, and email. It's on for everyone out of the box, and admins can raise the limit or switch it off.

Web research has its own limit. Dot can run about 10 web searches per question, and that allowance covers everything it does for that question, so a broad research request cannot quietly run up a bill.

You can also keep Dot on topic. Turn on **Business questions only** under Settings, then Advanced Settings, then AI & Data Configuration, and Dot politely declines personal or entertainment questions instead of working on them. It decides before running anything, so a declined question costs close to nothing. See [Business questions only](/using-dot/analyze#business-questions-only).

For anything more specific than that, write it as a note in Dot's context. A note is the right place for rules that are particular to your business, such as which topics to send to another team.


# Embed

enable your customers to chat with data & visualize insights

<div align="left"><figure><img src="/files/NfNuFvCkhC4D7gSWrTyv" alt=""><figcaption></figcaption></figure></div>

You can embed Dot directly in your web app via an Iframe and configure parts of the UI with the following parameters. They are all optional.

* **hideSideNavigation=true**
  * hides the complete left side navigation (default false)
* **hideHelp=true**
  * hides the little help dongle at the bottom right (default false)
* **hideShareButton=true**
  * hides the ability to share a conversation (default false)
* **hideTitle=true**
  * hides the title of a conversation (default false)
* **uiMode=dark**
  * specify either `dark` or `light` mode of the UI (default system)
* **hideExplanation=true**
  * hides full logs button & explanation tab
* **minimizeProgress=true**
  * only show 'thinking' as progess when waiting for an answer
* **hideNativeLinks=true**
  * will convert html links to javascript links, so users can't open a new tab accidentally
* **hideFileUpload=true**
  * hides the file upload button in the chat input (default false)
* **primaryColor=%23c700c7**
  * styles important action elements (e.g. chat input + button) according the specified color
  * the color is a hex code with the url encoded #c700c7
* **sideNavigationBackgroundColor=%23ffdbff**
  * styles the background color of the side navigation
* **chatInputBackgroundColor=%23ffdbff**
  * styles the background color of the chat input field
* **chatBackgroundColor=%23ffe6ff**
  * sets the background of the chat window
* **chatPlaceHolderText=Ask%20AI**
  * changes the placeholder text in the chat input field
  * default value is "Ask Dot about your data..."
* **suggestedPrompt1=How+is+the+weather**
  * changes the first suggested prompt text to the url-encoded value
* **suggestedPrompt\[2-8]=How+is+the+weather**
  * changes the nth suggested prompt text to the url-encoded value. You can specify up to 8 suggested questions on the chat window.
* **scope=MYDB%2EMYSCHEMA%2EMYTABLE**
  * specify the id of the data source that should be used as a scope for Dot to answer questions
  * the value is url encoded, e.g. MYDB.MYSCHEMA.MYTABLE
  * if this is not specified Dot will search for the right data source to answer the question
* **extra=anything**
  * here you can store anything that should get returned by the conversation history (e.g. an embedded user id)
* **active\_groups=finance%2Csales**
  * restricts a users data access to the specified user groups
  * Use `%2C` to url-encode a `,` to specify multiple groups
  * Active groups are validated against the assigned groups of the user

**Full example url**

```
https://eu.getdot.ai?hideSideNavigation=true&hideHelp=true&hideShareButton=true&hideExplanation=true&minimizeProgess=True&primaryColor=%23c700c7&chatPlaceHolderText=Ask%20AI
```

### Automatically login users

If you want to automatically login your users, you can pass the access\_token parameter into the url

* **access\_token=eyJhABCDE123456789IsInR5cCI6IkpXVCJ9.eyJzdWIiOiJrb2xhFGHIJ67890bGVkLnNvIiwib3JnX2secret2xlZC5zbyIsImV4cCI6MdotKL09876H0.spR-XrXTtDOTP54321ZWWchwR0x\_S8W\_juPVh8k**

The token can be obtained per user using the [api/auth/embedded\_user\_login ](/developers/api/commonly-used-endpoints#automatically-authenticate-embedded-users)endpoint.

Questions? Get [Support](/security-and-support/support)!


# API

automate as much as you like

## Everything starts with a token of trust

All API endpoints can be accessed via an API token that is tied to the permissions of a user account. You can also let them expire after some time.

### How to get a token?

1. Go to **Settings / Users**
2. Click **Create New Token**

<figure><img src="/files/vPdzqELbSVyFMEx2lO5e" alt=""><figcaption></figcaption></figure>

3. **Enter a name**, description, and expiration period
4. **Copy the token** (it's only shown once)

### How to use the token?

You have two ways. You either pass the token as a header with `API-KEY` or you pass it as a url parameter in `api_token` . As a header is usually more secure because automated loggers don't store them, but in some places you can't set headers (e.g. dbt webhooks) and then you can use the URL parameter.

#### Via Headers

Call the user endpoint via command line interface.

```bash
# Basic API request with token
curl -H "API-KEY: dot-your_token_here" <https://[app or eu].getdot.ai/api/auth/me>
```

Call the user endpoint via Python.

```python
import requests
headers = {"API-KEY": "dot-your_token_here"}
response = requests.get("https://[app or eu].getdot.ai/api/auth/me", headers=headers)
```

#### Via URL - Parameters

For the following endpoints you can also use a URL based authentication:

* Sync connection
* Import external asserts
* Export conversation history

Call the endpoint only via url:

```bash
curl "https://{region}.getdot.ai/api/sync/{connection_type}/{connection_id}" \
     "?user_id={user}&api_token={api_token}"
```

## Learn More

{% content-ref url="/pages/rUXw9XB23uXJ8TSwOjau" %}
[Commonly used Endpoints](/developers/api/commonly-used-endpoints)
{% endcontent-ref %}

{% content-ref url="/pages/NcBbF7phunv7RZRZqG7J" %}
[Use Cases and Scripts](/developers/api/use-cases-and-scripts)
{% endcontent-ref %}

{% content-ref url="/pages/3R5xLaAuLbOoxqCgtDPH" %}
[All API Endpoints](/developers/api/all-api-endpoints)
{% endcontent-ref %}


# Commonly used Endpoints

## Automatically Sync Dot

To keep Dot in sync with your production environment, it is recommended to trigger the following API endpoint

## POST /api/sync/{connection\_type}/{connection\_id}

> Sync connection with an API token.

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]},{"QueryTokenAuth":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}},"QueryTokenAuth":{"type":"apiKey","in":"query","name":"user_id+api_token"}},"schemas":{"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/sync/{connection_type}/{connection_id}":{"post":{"tags":["Connections"],"summary":"Sync connection with an API token.","operationId":"sync_connection_api_sync__connection_type___connection_id__post","parameters":[{"name":"connection_type","in":"path","required":true,"schema":{"type":"string","title":"Connection Type"}},{"name":"connection_id","in":"path","required":true,"schema":{"type":"string","title":"Connection Id"}}],"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

```javascript
// URL endpoint
https://{region}.getdot.ai/api/sync/{connection_type}/{connection_id}?user_id={user}&api_token={api_token}
```

* **Region**: `app` (US) or `eu` (EU)
* **Connection Type**: `postgres`, `redshift`, `snowflake`, `mssql`, `bigquery`, `databricks`, `looker`, `dbt`
* **User ID**: usually email of the user [(url encoded)](https://www.urlencoder.io/)
* **API Token**: can get created (and overwritten) by clicking `Copy API Token` in Settings/Users/Actions/···

\
**Trigger with curl (CLI)**

```javascript
curl -X "POST" "https://eu.getdot.ai/api/sync/bigquery/my-bg-id?user_id=sync_user%40contoso.com&api_token=dot-42673584be9724a21e1550336d6fe509f4a04207461ec9a926ca2a27cbd27fa0
```

**Trigger with dbt webhooks**

Call the api endpoint after your dbt run completed.

{% embed url="<https://docs.getdbt.com/docs/deploy/webhooks#create-a-webhook-subscription>" %}
Documentation how to setup a dbt webhooks
{% endembed %}

## Import External Assets

Inform Dot about key external knowledge assets, such as BI dashboards or custom data apps, so it can recommend them to users and assist with discovery and understanding. Authentication works similarly to the Sync Connection endpoint.

## POST /api/import\_and\_overwrite\_external\_asset

> Import and overwrite an external asset with an API token.

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]},{"QueryTokenAuth":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}},"QueryTokenAuth":{"type":"apiKey","in":"query","name":"user_id+api_token"}},"schemas":{"Body_import_external_asset_api_import_and_overwrite_external_asset_post":{"properties":{"external_asset":{"$ref":"#/components/schemas/ExternalAsset"}},"type":"object","required":["external_asset"],"title":"Body_import_external_asset_api_import_and_overwrite_external_asset_post"},"ExternalAsset":{"properties":{"id":{"type":"string","title":"Id","description":"Unique identifier for the external asset."},"subtype":{"type":"string","title":"Subtype","description":"Subtype of the external asset, such as 'looker_dashboard'."},"folder":{"anyOf":[{"type":"string"},{"type":"null"}],"title":"Folder","description":"Folder path where the external asset is stored."},"name":{"anyOf":[{"type":"string"},{"type":"null"}],"title":"Name","description":"The name of the external asset."},"description":{"anyOf":[{"type":"string"},{"type":"null"}],"title":"Description","description":"Description of the external asset."},"url":{"anyOf":[{"type":"string"},{"type":"null"}],"title":"Url","description":"Full URL of the external asset."},"active":{"type":"boolean","title":"Active","description":"Indicates if the external asset is active.","default":true},"view_count":{"anyOf":[{"type":"integer"},{"type":"null"}],"title":"View Count","description":"Number of times the external asset has been viewed.","default":0},"last_accessed_at":{"anyOf":[{"type":"string"},{"type":"null"}],"title":"Last Accessed At","description":"Timestamp of the last access to the external asset."},"dot_description":{"anyOf":[{"type":"string"},{"type":"null"}],"title":"Dot Description","description":"Formatted text description of the external asset, can e.g. include a markdown representation of all charts in a dashboard."}},"type":"object","required":["id","subtype"],"title":"ExternalAsset"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/import_and_overwrite_external_asset":{"post":{"tags":["Assets"],"summary":"Import and overwrite an external asset with an API token.","operationId":"import_external_asset_api_import_and_overwrite_external_asset_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_import_external_asset_api_import_and_overwrite_external_asset_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

## Export Conversation History

Export all conversations together with relevant meta data fields such as number of messages or author.

## GET /api/export\_history

> Get all historical chat messages.

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]},{"QueryTokenAuth":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}},"QueryTokenAuth":{"type":"apiKey","in":"query","name":"user_id+api_token"}},"schemas":{"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/export_history":{"get":{"tags":["Admin"],"summary":"Get all historical chat messages.","operationId":"export_history_api_export_history_get","parameters":[{"name":"start_date","in":"query","required":false,"schema":{"type":"string","description":"by default two weeks ago, format YYYY-MM-DD","title":"Start Date"},"description":"by default two weeks ago, format YYYY-MM-DD"},{"name":"end_date","in":"query","required":false,"schema":{"type":"string","description":"by default tomorrow, format YYYY-MM-DD","title":"End Date"},"description":"by default tomorrow, format YYYY-MM-DD"},{"name":"include_all_workspaces","in":"query","required":false,"schema":{"type":"boolean","description":"Include history from all workspaces (main workspace only)","default":false,"title":"Include All Workspaces"},"description":"Include history from all workspaces (main workspace only)"}],"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

## Ask Dot Automatically

`/api/agentic` is how you ask Dot anything. It is the same engine behind the app, Slack and Teams: Dot reads your model, writes and runs SQL, builds charts, and keeps going until it can answer. A one-line lookup and a multi-step investigation use this one endpoint — Dot decides how far to go.

Send a question with a `chat_id` you generate. The call returns the conversation once Dot has finished.

## POST /api/agentic

> Agentic

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}},"schemas":{"Body_agentic_api_agentic_post":{"properties":{"messages":{"items":{"additionalProperties":true,"type":"object"},"type":"array","title":"Messages"},"chat_id":{"type":"string","title":"Chat Id"},"skip_check":{"type":"boolean","title":"Skip Check","default":false},"extra":{"anyOf":[{"type":"string"},{"type":"null"}],"title":"Extra"}},"type":"object","required":["messages","chat_id"],"title":"Body_agentic_api_agentic_post"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/agentic":{"post":{"tags":["Chat"],"summary":"Agentic","operationId":"agentic_api_agentic_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_agentic_api_agentic_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

Follow up by posting to `/api/agentic_with_history` with the same `chat_id`, so Dot keeps the context of everything already asked.

Both take an optional `mode` — `economy`, `balanced` or `frontier` — the same dial as the [energy mode](/using-dot/analyze#energy-modes) selector in the app. Left out, your workspace default applies.

To ask against an [environment](/train-dot/environments) instead of production, send its id in the `X-Dot-Environment` header. That works on every endpoint on this page, not just these two.

The answer sits in the last message carrying a `formatted_result`: a list of parts, of which the `text` ones joined together are what a person reads. A long investigation can outlast the request, so fetch it later with `GET /api/c2/{chat_id}` using the same `chat_id`. There is a [worked example](/developers/api/use-cases-and-scripts) of both.

## User Administration

## GET /api/get\_users

> Get All Users

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}}},"paths":{"/api/get_users":{"get":{"tags":["Admin"],"summary":"Get All Users","operationId":"get_all_users_api_get_users_get","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}}}}
```

## POST /api/send\_invitations

> Sendinvitations

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}},"schemas":{"Body_sendInvitations_api_send_invitations_post":{"properties":{"emails":{"items":{"type":"string"},"type":"array","title":"Emails"}},"type":"object","required":["emails"],"title":"Body_sendInvitations_api_send_invitations_post"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/send_invitations":{"post":{"tags":["Admin"],"summary":"Sendinvitations","operationId":"sendInvitations_api_send_invitations_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_sendInvitations_api_send_invitations_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

## POST /api/delete\_user

> Delete User Route

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}},"schemas":{"Body_delete_user_route_api_delete_user_post":{"properties":{"email":{"type":"string","title":"Email"}},"type":"object","required":["email"],"title":"Body_delete_user_route_api_delete_user_post"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/delete_user":{"post":{"tags":["Admin"],"summary":"Delete User Route","operationId":"delete_user_route_api_delete_user_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_delete_user_route_api_delete_user_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

## POST /api/change\_user\_role

> Change User Role

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}},"schemas":{"Body_change_user_role_api_change_user_role_post":{"properties":{"email":{"type":"string","title":"Email"},"role":{"type":"string","title":"Role"}},"type":"object","required":["email","role"],"title":"Body_change_user_role_api_change_user_role_post"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/change_user_role":{"post":{"tags":["Admin"],"summary":"Change User Role","operationId":"change_user_role_api_change_user_role_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_change_user_role_api_change_user_role_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

## POST /api/add\_user\_to\_group

> Add User To Group

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}},"schemas":{"Body_add_user_to_group_api_add_user_to_group_post":{"properties":{"email":{"type":"string","title":"Email"},"group":{"type":"string","title":"Group"}},"type":"object","required":["email","group"],"title":"Body_add_user_to_group_api_add_user_to_group_post"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/add_user_to_group":{"post":{"tags":["Admin"],"summary":"Add User To Group","operationId":"add_user_to_group_api_add_user_to_group_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_add_user_to_group_api_add_user_to_group_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

## POST /api/remove\_user\_from\_group

> Remove User From Group

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}},"schemas":{"Body_remove_user_from_group_api_remove_user_from_group_post":{"properties":{"email":{"type":"string","title":"Email"},"group":{"type":"string","title":"Group"}},"type":"object","required":["email","group"],"title":"Body_remove_user_from_group_api_remove_user_from_group_post"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/remove_user_from_group":{"post":{"tags":["Admin"],"summary":"Remove User From Group","operationId":"remove_user_from_group_api_remove_user_from_group_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_remove_user_from_group_api_remove_user_from_group_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

## POST /api/create\_user

> Create User Route

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}},"schemas":{"Body_create_user_route_api_create_user_post":{"properties":{"email":{"type":"string","title":"Email"},"password":{"type":"string","title":"Password"},"realname":{"type":"string","title":"Realname"}},"type":"object","required":["email","password","realname"],"title":"Body_create_user_route_api_create_user_post"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/create_user":{"post":{"tags":["Admin"],"summary":"Create User Route","operationId":"create_user_route_api_create_user_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_create_user_route_api_create_user_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```

## Automatically Authenticate Embedded Users

For embedded use cases that require SSO, where your end users have individual permissions you can use this endpoint to obtain an access token for users that is valid for 24h. Here is an example on how you can use it to [embed](/developers/embed) Dot in your application.

Please make sure that you enabled this flag on settings: **"Allow admins to authenticate for users to enable SSO in embeds".**

## POST /api/auth/embedded\_user\_login

> Get Embedded User Login Token

```json
{"openapi":"3.1.0","info":{"title":"Dot API","version":"1.0.0"},"security":[{"APIKeyHeader":[]},{"APIKeyHeader":[]},{"APIKeyCookie":[]},{"OAuth2PasswordBearer":[]}],"components":{"securitySchemes":{"APIKeyHeader":{"type":"apiKey","in":"header","name":"X-API-KEY"},"APIKeyCookie":{"type":"apiKey","in":"cookie","name":"Authorization"},"OAuth2PasswordBearer":{"type":"oauth2","flows":{"password":{"scopes":{},"tokenUrl":"/api/auth/token"}}}},"schemas":{"Body_get_embedded_user_login_token_api_auth_embedded_user_login_post":{"properties":{"user_id":{"type":"string","title":"User Id"},"create_if_not_exists":{"type":"boolean","title":"Create If Not Exists","default":false},"groups":{"anyOf":[{"items":{"type":"string"},"type":"array"},{"type":"null"}],"title":"Groups"}},"type":"object","required":["user_id"],"title":"Body_get_embedded_user_login_token_api_auth_embedded_user_login_post"},"HTTPValidationError":{"properties":{"detail":{"items":{"$ref":"#/components/schemas/ValidationError"},"type":"array","title":"Detail"}},"type":"object","title":"HTTPValidationError"},"ValidationError":{"properties":{"loc":{"items":{"anyOf":[{"type":"string"},{"type":"integer"}]},"type":"array","title":"Location"},"msg":{"type":"string","title":"Message"},"type":{"type":"string","title":"Error Type"}},"type":"object","required":["loc","msg","type"],"title":"ValidationError"}}},"paths":{"/api/auth/embedded_user_login":{"post":{"tags":["Authentication"],"summary":"Get Embedded User Login Token","operationId":"get_embedded_user_login_token_api_auth_embedded_user_login_post","requestBody":{"content":{"application/json":{"schema":{"$ref":"#/components/schemas/Body_get_embedded_user_login_token_api_auth_embedded_user_login_post"}}},"required":true},"responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}},"422":{"description":"Validation Error","content":{"application/json":{"schema":{"$ref":"#/components/schemas/HTTPValidationError"}}}}}}}}}
```


# Use Cases and Scripts

Here is a list of common ways to use the API. All examples are done in Python.

Prerequisite: [Get an API token](/developers/api#everything-starts-with-a-token-of-trust).

<details>

<summary>Add labels to a conversation</summary>

```python
import requests

# Configuration
BASE_URL = "https://app.getdot.ai"
ADD_LABEL_ENDPOINT = "/api/add_label_to_chat"

# Replace with your API token obtained from your account settings.
API_TOKEN = "dot-your_token_here"

# Chat details
CHAT_ID = "your_chat_id"
LABELS = ["your_label"]

def add_label_to_chat(chat_id, labels):
    """Add labels to a chat using token-based authentication."""
    headers = {
        "API-KEY": API_TOKEN,
        "Content-Type": "application/json"
    }
    data = {"chat_id": chat_id, "labels": labels}
    
    try:
        response = requests.post(f"{BASE_URL}{ADD_LABEL_ENDPOINT}", headers=headers, json=data)
        response.raise_for_status()
        print("Successfully added labels to chat.")
    except requests.exceptions.RequestException as e:
        print(f"Failed to add labels to chat: {e}")

def main():
    add_label_to_chat(CHAT_ID, LABELS)

if __name__ == "__main__":
    main()
```

</details>

<details>

<summary>Ask a question to Dot and follow up</summary>

```python
"""
Dot API Client Example

A minimal example showing how to interact with the Dot API to ask questions
about your data and follow up with additional questions in the same conversation.

Posting a question returns the conversation once Dot has finished answering.
A long investigation can outlast the request, so the example also shows how to
poll for the answer afterwards using the same chat_id.

This pattern applies to both initial questions and follow-up questions.

Usage:
    python3 test_api.py

Requirements:
    - Python 3.6+
    - requests library (pip install requests)
"""

import requests
import uuid

# Replace with your Dot API key from the Settings page
API_KEY = "dot-YOUR_API_KEY_HERE"

# Use "https://eu.getdot.ai/api" for the EU region
BASE_URL = "https://app.getdot.ai/api"
HEADERS = {"API-KEY": API_KEY, "Content-Type": "application/json"}


def ask_question(question):
    """
    Send a question to Dot and return the conversation.
    
    Returns:
        tuple: (messages, chat_id)
    """
    chat_id = str(uuid.uuid4())
    
    print(f"Asking question: '{question}'")
    endpoint = f"{BASE_URL}/agentic"
    payload = {"messages": [{"role": "user", "content": question}], "chat_id": chat_id}
    # Add "mode": "economy" | "balanced" | "frontier" to control how deep Dot goes
    
    response = requests.post(endpoint, headers=HEADERS, json=payload, timeout=600)
    response.raise_for_status()
    
    return response.json(), chat_id


def ask_follow_up(question, chat_id):
    """
    Send a follow-up question using the same chat session.
    
    Returns:
        list: The updated conversation
    """
    print(f"Asking follow-up: '{question}'")
    endpoint = f"{BASE_URL}/agentic_with_history"
    payload = {"new_message": {"role": "user", "content": question}, "chat_id": chat_id}
    
    response = requests.post(endpoint, headers=HEADERS, json=payload, timeout=600)
    response.raise_for_status()
    
    return response.json()


def fetch_conversation(chat_id):
    """
    Read a conversation back later, or poll for an answer that outlasted the request.
    
    Note the shape difference: posting a question returns the messages as a list,
    while this endpoint wraps them in {"messages": [...]}.
    """
    response = requests.get(f"{BASE_URL}/c2/{chat_id}", headers=HEADERS)
    response.raise_for_status()
    return response.json().get("messages", [])


def answer_text(messages):
    """
    Pull the user-visible answer out of a conversation.
    
    Dot's answer is the last message carrying a formatted result. That result is a
    list of parts — text, tables, charts — and the text parts joined together are
    what a person reads in the app.
    """
    if isinstance(messages, dict):
        messages = messages.get("messages", [])
    
    for message in reversed(messages or []):
        parts = (message.get("additional_data") or {}).get("formatted_result") or []
        texts = [str(p["data"]) for p in parts if p.get("type") == "text" and p.get("data")]
        if texts:
            return "\n\n".join(texts)
        
        content = message.get("content")
        if isinstance(content, str) and content.strip() and content.strip().lower() != "success":
            return content.strip()
    
    return ""


def print_response(messages):
    answer = answer_text(messages)
    if answer:
        print("\n=== ANSWER ===")
        print(answer)
        print("\n")
    else:
        print("No answer yet — try fetch_conversation(chat_id) in a few seconds.")


def main():
    """A question, then a follow-up in the same conversation."""
    try:
        initial_question = "What were our total sales last month?"
        try:
            user_input = input("Enter your question: ")
            if user_input.strip():
                initial_question = user_input
        except EOFError:
            print(f"Using default question: '{initial_question}'")
        
        messages, chat_id = ask_question(initial_question)
        print_response(messages)
        print(f"Chat ID: {chat_id} (save this if you want to continue the conversation later)")
        
        follow_up = "How does that compare to the previous month?"
        try:
            user_input = input("Enter a follow-up question: ")
            if user_input.strip():
                follow_up = user_input
        except EOFError:
            print(f"Using default follow-up: '{follow_up}'")
        
        follow_up_response = ask_follow_up(follow_up, chat_id)
        print_response(follow_up_response)
        
        print("Conversation complete! You can continue by using the same chat_id.")
        
    except Exception as e:
        print(f"Error: {e}")
        import traceback
        traceback.print_exc()


if __name__ == "__main__":
    main()
```

</details>

<details>

<summary>Sync a confluence page to the note</summary>

```python
import os, re, requests, markdownify

# ---------- 1)  Confluence ----------------------------------------------------
ATL_SITE  = "https://<your-site>.atlassian.net/wiki"
PAGE_ID   = "<confluence_page_id>"
ATL_AUTH  = (os.getenv("ATLASSIAN_EMAIL"), os.getenv("ATLASSIAN_API_TOKEN"))

r = requests.get(
    f"{ATL_SITE}/rest/api/content/{PAGE_ID}?expand=body.storage",
    auth=ATL_AUTH,
)
r.raise_for_status()
html = r.json()["body"]["storage"]["value"]
md   = markdownify.markdownify(html, heading_style="ATX")
page_url = f"{ATL_SITE}/pages/{PAGE_ID}"

# ---------- 2)  Dot – read org & note ----------------------------------------
DOT_BASE = "https://eu.getdot.ai/api"   # or https://app.getdot.ai/api for US
HEADERS  = {"API-KEY": os.getenv("DOT_API_KEY")}

notes    = requests.get(f"{DOT_BASE}/org_notes", headers=HEADERS).json()
existing = next((n for n in notes if n.get("title") == "Confluence FAQ"), None)
note     = existing.get("note", "") if existing else ""

# ---------- 3)  insert / replace the <faq> block ------------------------------
new_faq = f'<faq confluence_page_url="{page_url}">\n\n{md}\n\n</faq>'
note    = re.sub(r"<faq[^>]*>.*?</faq>", new_faq, note, flags=re.I|re.S) \
          if "<faq" in note.lower() else f"{note.rstrip()}\n\n{new_faq}"

# ---------- 4)  Dot – save the updated note -----------------------------------
if existing:
    requests.put(f"{DOT_BASE}/org_notes/{existing['id']}", headers=HEADERS,
                 json={"title": "Confluence FAQ", "note": note}).raise_for_status()
else:
    requests.post(f"{DOT_BASE}/org_notes", headers=HEADERS,
                  json={"title": "Confluence FAQ", "note": note}).raise_for_status()

print("✅  org‑note updated")
```

</details>


# All API Endpoints

Once you [created your token](/developers/api#everything-starts-with-a-token-of-trust), you can use [**all API endpoints**](https://test.getdot.ai/redoc).


# CLI & AI Agent Skill

Give your AI coding assistant direct access to your company's data. One install, and Claude Code, Cursor, Codex, and Gemini CLI can query your databases while you work.

Most data questions interrupt your flow. You leave your editor, open a browser, wait for an answer, copy it back. The Dot skill eliminates that loop.

Once installed, Claude Code treats Dot like a sub-agent. It delegates data questions to Dot automatically: "what were our top customers last quarter?" gets answered in seconds, right in your terminal. You never context-switch.

<figure><img src="/files/0pqvpo6KUFCmBRWWeifW" alt="Claude Code loading the Dot skill via /dot command"><figcaption><p>Type /dot in Claude Code to query your company data</p></figcaption></figure>

Works with **Claude Code**, **Cursor**, **OpenAI Codex**, and **Gemini CLI**. The installer detects which agents you have and configures the skill for each one.

### Install

Use the URL that matches your Dot region:

| Region | URL             |
| ------ | --------------- |
| US     | `app.getdot.ai` |
| EU     | `eu.getdot.ai`  |

**macOS / Linux:**

```bash
# US region
curl https://app.getdot.ai/install | sh

# EU region
curl https://eu.getdot.ai/install | sh
```

**Windows (PowerShell):**

```powershell
# US region
irm https://app.getdot.ai/install.ps1 | iex

# EU region
irm https://eu.getdot.ai/install.ps1 | iex
```

Then log in:

```bash
dot login
```

Opens your browser to authenticate. Token is saved locally.

That's it. Your AI agent can now query your data.

{% hint style="info" %}
You can also install from the **Set Up CLI** page in your Dot dashboard (`/cli-setup`). It generates a command with your auth token embedded so you skip the login step.
{% endhint %}

**Self-hosted Dot:**

```bash
curl https://your-dot-instance.com/install | sh
```

**CI / headless servers:**

```bash
dot login --token <YOUR_API_TOKEN>
```

### How it works

The installer puts a native `dot` binary on your PATH and drops a `SKILL.md` file into the right location for each AI agent. That skill file teaches your coding assistant when and how to call Dot. It's the bridge between a natural language question and your data.

When an agent invokes the skill, it runs `dot` as a sub-agent. Dot queries your database and returns a structured answer: text explanation, data preview, chart, CSV, and a link to the full analysis in the browser. The agent reads all of this and uses it in context.

No runtime dependencies. No configuration beyond `dot login`.

### Using with AI agents

#### Claude Code

Type **`/dot`** followed by your question:

```
/dot What data sources do we have?
```

Or just ask naturally. Claude Code will delegate to Dot when it recognizes a data question:

* "What were our top 10 customers by revenue last quarter?"
* "Is the orders table growing?"
* "Show me monthly active users for the past year"

Claude Code uses Dot like any other tool in its toolkit. It decides when to call it, reads the results, and incorporates them into its response.

#### Cursor, Codex, and Gemini CLI

Same behavior. The installer configures each agent it detects. The skill file tells the agent what Dot can do and when to use it. No manual setup.

### Standalone usage

You can also use Dot directly from any terminal:

```bash
dot "What were total sales last month?"
```

Returns a text explanation, data preview, chart (PNG), CSV data, a link to the full analysis in Dot, and suggested follow-ups.

#### Follow-up questions

Every response includes a chat ID. Continue the conversation:

```bash
dot "Now break down by region" --chat cli-m1abc2d-x4y5z6
```

#### View available data

```bash
dot catalog
```

Instant response. Shows your connections, tables, column counts, row counts, and any external assets like Looker dashboards.

#### Other commands

```bash
dot update          # Update to latest version
dot status          # Login status and token info
dot logout          # Clear credentials
dot --version       # Show version
dot --help          # Show all options
```

### Workspace management

The CLI is designed as the AI interface to Dot. When you use Claude Code, Cursor, or Codex, your AI assistant manages your Dot workspace through these commands — selecting relevant tables, creating notes with metric definitions, and fixing context gaps. You don't need to type these commands yourself; your AI agent handles it.

{% hint style="info" %}
**How AI agents use this:** When you say "help me set up Dot for our revenue data", your AI assistant will run `dot catalog` to see what's available, `dot tables list` to check active tables, `dot notes list` to audit existing context, and then interview you to fill gaps — creating notes from your answers.
{% endhint %}

#### Notes

Notes teach Dot how to query your data. They contain metric definitions, SQL patterns, business context, and operating principles. See [Notes](https://github.com/Snowboard-Software/snowboard_software/tree/main/dot/whats-dot/model/notes.md) for what to put in them.

```bash
dot notes list                          # List all notes (tree view)
dot notes get <id>                      # Read a note
dot notes create <title> [content]      # Create a note
dot notes update <id> [content]         # Update a note
dot notes delete <id>                   # Delete a note
dot notes activate <id>                 # Activate a note
dot notes deactivate <id>              # Deactivate a note
```

Create notes from files or stdin for longer content:

```bash
dot notes create "Metric Glossary" --file metrics.md
cat playbook.md | dot notes create "Campaign Review Playbook"
```

Organize notes with groups and parent relationships:

```bash
dot notes create "SQL Patterns" --parent <parent-id> --group data-team
```

Inactive notes are not used by Dot when answering questions. Deactivate instead of deleting to archive notes you might need later.

#### Tables

Control which database tables Dot can access. Deactivated tables are hidden from Dot entirely.

```bash
dot tables list                         # Active tables, grouped by connection
dot tables list --all                   # Include inactive tables
dot tables get <id>                     # Table details with columns
dot tables activate <id> [id2 ...]      # Activate table(s)
dot tables deactivate <id> [id2 ...]    # Deactivate table(s)
```

Batch operations work with multiple IDs:

```bash
dot tables activate table-1 table-2 table-3
```

#### Chat history

Retrieve the full conversation history for any previous query:

```bash
dot chat <chat-id>
```

Shows all messages with SQL queries, tool calls, and timestamps. Useful for auditing what Dot did or retrieving answers you saw earlier.

#### Feature requests

Submit feedback or feature requests directly from the terminal:

```bash
dot wish "I want better date filtering"
```

### Apps as code

Dashboards ("apps") are code too. Author `.app` files locally, push them to Dot to build and activate, and manage them through your GitHub repo like any other source. Pushed apps land in the `apps/` folder that [GitHub Sync](https://github.com/Snowboard-Software/snowboard_software/tree/main/dot/whats-dot/version-control/github.md) mirrors to your repo.

The loop:

1. **Author** or edit an `.app` file locally.
2. **`dot apps push <id>`** — Dot builds and activates the app, then writes an `.app.lock` (the compiled SQL) next to it.
3. **Sync** — your `apps/` folder pushes to your connected GitHub repo.
4. **`dot apps pr`** — commit the app files and open a pull request.

Common flow:

```bash
dot apps pull <id>       # Fetch the .app source + .app.lock
dot apps push <id>       # Build + activate; writes the .app.lock
dot apps status <id>     # Lock freshness (non-zero exit when stale)
dot apps preview <id>    # Print the preview URL
dot apps pr              # Commit app files, push, open a PR
```

`dot apps deploy` is an alias for `dot apps pr`.

Advanced:

```bash
dot apps validate <id>                 # Validate locks against the model
dot apps bless <id>                    # Accept hand-edited compiled SQL
dot apps resolve <id>                  # Recompile stale queries
dot apps refactor <id> --rename old:new  # Rewrite locked SQL (sqlglot)
dot apps deps <table[.column]>         # Show apps that depend on a table/column
dot apps plan                          # Derive dbt schema mapping and renames
```

{% hint style="info" %}
`dot apps status` exits non-zero when a lock is stale, so you can gate CI on it. Add `--json` for machine-readable output.
{% endhint %}

### Environments

Environments let you branch and test model changes safely before they reach Production — each one is a git branch of your model that you can edit, diff, and merge back.

```bash
dot env list             # Production + all environments
dot env current          # The active environment
dot env show <name>      # Environment details
dot env create <name>    # New environment (--from <id|name> to branch off another)
dot env use <name>       # Set the active environment (scopes later commands)
dot env diff <name>      # Changed files vs Production
dot env merge <name> --confirm   # Merge changes back into Production
```

`dot env use` saves your choice locally; every later command runs against that environment until you `dot env unset`. Further verbs — `rename`, `delete`, `unset`, `conflicts` (predict merge conflicts), and `target` / `sync-target` (per-connection dbt-style overrides) — are covered by:

```bash
dot env --help
```

### Training Dot effectively

The biggest leverage in Dot accuracy comes from good context management. Here's the recommended approach:

**Start with tables.** Select the right data sources and document what each table represents — especially granularity (what does one row mean?) and primary keys. For wide tables or semantic models with hundreds of fields, curate which columns are active.

**Layer in notes.** After tables are set up, add notes for metric definitions, operating principles, and business context. Use strong, unambiguous language: "Always filter out internal users", "Revenue means recognized revenue, not bookings."

**Fix bad answers at the root.** When Dot gives a wrong answer, don't just correct it — investigate why. Wrong table? Activate the right one. Wrong logic? Add a metric definition. Missing context? Add a business context note.

**Audit regularly.** Review existing notes for contradictions. If one note says "use fct\_orders for revenue" and another says "use fct\_arr for revenue" — that ambiguity will produce inconsistent answers. Remove or reconcile.

**Front-load, then prune.** It's better to give too much context upfront and dial it back than to discover gaps one bad answer at a time.

### Caching

Responses are cached on disk. Repeated questions return instantly.

```bash
dot "question" --no-cache    # Skip cache, force fresh
dot --clear-cache            # Clear all cached responses
```

Follow-ups (`--chat`) and `dot catalog` are never cached. Cache lives at `~/.cache/dot/`, scoped per user.

### Security

* Tokens stored locally with `600` permissions
* One token per user. Generating a new one revokes the old one
* Tokens expire after 365 days
* All Dot permissions and row-level security enforced
* All queries logged for compliance

### Troubleshooting

**"Not authenticated"** Run `dot login` or check with `dot status`.

**"Connection failed"** Check your network. For custom servers: `dot login --server https://your-server.com`.

**Slow responses** First query takes 10-30 seconds (full analysis pipeline). Identical queries return instantly from cache after that.


# MCP

### What is MCP?

MCP (Model Context Protocol) lets AI assistants like Claude, Cursor, and ChatGPT connect directly to your data sources through secure connections. With Dot's MCP integration, your assistant can query your business data while respecting all of your existing Dot permissions and setup.

### Your Dot MCP URL

You'll paste this URL into every client below. It depends on which region you sign into Dot from:

* **US**: `https://app.getdot.ai/ai/mcp`
* **EU**: `https://eu.getdot.ai/ai/mcp`

Use the same host you use to sign into Dot in the browser. You can also copy the exact URL from **Settings → Integrations** inside Dot.

{% hint style="success" %}
**OAuth is the default.** For supported clients (Claude, Cursor, Windsurf), OAuth means no tokens to copy or manage — just sign into Dot in your browser and click **Allow**. Sessions last up to a year.

If OAuth isn't supported by your client (ChatGPT, Raycast, generic MCP clients) or something goes wrong, see [Using an API token](#using-an-api-token).
{% endhint %}

### Claude (Web, Desktop, Cowork, Mobile)

**Requirements**: Free, Pro, Max, Team, or Enterprise plan. Free users can connect one custom connector at a time.

{% hint style="info" %}
**You only need to set this up once.** Custom connectors are remote MCP servers hosted by Dot, so Anthropic syncs them across every Claude surface automatically. Adding Dot on claude.ai makes it available in Claude Desktop, Cowork, and the Claude mobile apps with no per-device setup. Prefer to start from Claude Desktop? Open **Settings → Connectors → Add custom connector** and follow the same steps — it'll sync back to the web.
{% endhint %}

{% tabs %}
{% tab title="Pro / Max" %}
{% stepper %}
{% step %}

### Open Customize → Connectors

In Claude, click your profile → **Customize** → **Connectors**.

<figure><img src="/files/CBNelGMBmEoZs8w84NVf" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/VsytsUwebOSVqtvPdo2V" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

### Add a custom connector

Click the **+** next to the Connectors header and choose **Add custom connector**.

<figure><img src="/files/KXOLmh9AicecItZqX3ro" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

### Name it and paste the URL

Name it **Ask Dot**, paste your Dot MCP URL, then click **Add**.

<figure><img src="/files/S1dMAnnY6ocQLHlBWIDe" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

### Connect

Click **Connect** on the new **Ask Dot** entry.

<figure><img src="/files/Pp9VmanQqyf0JA1rNsbp" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

### Authorize in Dot

A Dot tab opens. Review the requested permissions and click **Allow**.

<figure><img src="/files/sQ9uF2KxdJBfLHKKpbEm" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

### Done

Back in Claude, **Ask Dot** now shows as **Connected** with its available tools. Enable it in any chat via the **+** button → **Connectors**.

<figure><img src="/files/cUREXG0dnWiKXeaqTOW4" alt=""><figcaption></figcaption></figure>
{% endstep %}
{% endstepper %}
{% endtab %}

{% tab title="Team / Enterprise" %}
Team and Enterprise workspaces use a two-step flow: an owner adds the connector once for the organization, then each member connects their own Dot account.

**Owner (one-time setup):**

{% stepper %}
{% step %}

### Open Organization settings → Connectors

In Claude, go to **Organization settings → Connectors**.
{% endstep %}

{% step %}

### Add a custom web connector

Click **Add** → **Custom** → **Web**.
{% endstep %}

{% step %}

### Paste the Dot MCP URL

Name it **Ask Dot**, paste your Dot MCP URL, then click **Add**.
{% endstep %}
{% endstepper %}

**Each member (once):**

{% stepper %}
{% step %}

### Open Customize → Connectors

Find **Ask Dot** in the list — it'll be marked **Custom**.
{% endstep %}

{% step %}

### Connect

Click **Connect**.

<figure><img src="/files/Pp9VmanQqyf0JA1rNsbp" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

### Authorize in Dot

A Dot tab opens. Review the permissions and click **Allow**.

<figure><img src="/files/sQ9uF2KxdJBfLHKKpbEm" alt=""><figcaption></figcaption></figure>
{% endstep %}
{% endstepper %}
{% endtab %}
{% endtabs %}

### Claude Code

Run the command from Dot's Integrations page:

```bash
claude mcp add --transport http ask_dot https://app.getdot.ai/ai/mcp
```

OAuth will open your browser to authenticate on first use.

{% hint style="info" %}
Claude Code with a signed-in Anthropic account also picks up custom connectors synced from your Claude Web/Desktop setup. If you've already connected Dot there, you can skip this command.
{% endhint %}

### Cursor IDE

The recommended setup uses the [`mcp-remote`](https://www.npmjs.com/package/mcp-remote) stdio proxy — the same approach Linear, Cloudflare, and Sentry recommend for their MCP servers. **Requires** [**Node.js**](https://nodejs.org) **installed locally.**

**One-click install:** Click the **Add to Cursor** button on Dot's Integrations page. Cursor will open your browser to authorize.

**Manual install:** Add this to `~/.cursor/mcp.json` (use `https://eu.getdot.ai/ai/mcp` for the EU instance):

```json
{
  "mcpServers": {
    "ask_dot": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://app.getdot.ai/ai/mcp"]
    }
  }
}
```

Save the file, then restart Cursor (or toggle the server in **Settings → Tools & Integrations → MCP Tools**). On first connect, Cursor will spawn `mcp-remote`, open your browser to authorize, and cache the session for up to a year.

{% hint style="info" %}
If you'd rather use a static API token instead of OAuth, see [Using an API token → JSON config](#using-an-api-token) below for the Cursor-specific shape.
{% endhint %}

### Windsurf

Copy the URL from Dot's Integrations page and add it as an MCP server in Windsurf settings. Windsurf will open your browser to authenticate.

### ChatGPT

**Requirements**: ChatGPT Enterprise, Education, or Team subscription.

ChatGPT supports token-based auth only — see [Using an API token](#using-an-api-token) for how to generate a token. Then in ChatGPT's connector settings, add a new connector:

* **Name**: Ask\_dot
* **Description**: Dot AI-powered data analysis platform. \<Add additional info about the data you have connected and when to use it>
* **MCP Server URL**: the URL from Dot's Integrations page
* **Authentication**: your MCP token, sent in an `X-API-KEY` header (see [Using an API token](#using-an-api-token))

### Raycast AI

Raycast uses token-based auth — see [Using an API token](#using-an-api-token) to generate one.

1. Run **"Manage MCP Servers"** command in Raycast
2. Press `Cmd + N` to add a new server
3. Paste the configuration from Dot's Integrations page
4. Submit and use by @-mentioning "dot" in Raycast AI

<figure><img src="/files/81QyQl9hmOoQ5CzvuNbA" alt=""><figcaption></figcaption></figure>

### Other MCP clients

Most MCP clients support URL-based configuration. If your client supports OAuth / MCP authorization, use the plain Dot MCP URL — that's it. Otherwise, follow [Using an API token](#using-an-api-token) and send the token in a header.

### Using an API token

Use a token when your client doesn't support OAuth (ChatGPT, Raycast, generic MCP clients) or as a fallback if OAuth isn't working.

#### Generate a token

1. Go to **Settings → Integrations** in Dot
2. Select your client and switch to the **API Token** tab
3. Click **Generate MCP Token** and copy it immediately — you won't see it again

#### Apply the token

Depending on how your client accepts credentials, use one of these:

{% tabs %}
{% tab title="Header" %}
Send the token in an `X-API-KEY` header. That's how token auth works now, so use it anywhere your client lets you set request headers.

For Claude Code, pass the header on the add command:

```bash
claude mcp add --transport http ask_dot https://app.getdot.ai/ai/mcp --header "X-API-KEY: YOUR_TOKEN"
```

If your client only takes a plain URL and can't set headers, use OAuth instead (see the top of this page). Putting the token in the URL is no longer supported.
{% endtab %}

{% tab title="JSON config (Cursor, etc.)" %}
For clients that accept an MCP server block:

1. In Cursor, go to **Settings → Tools & Integration → Add new MCP Server** to open `mcp.json`
2. Grab the URL and API key from the JSON config section under "Others" on Dot's Integrations page
3. Add to `mcpServers`:

```json
{
  "mcpServers": {
    "ask_dot": {
      "url": "https://app.getdot.ai/ai/mcp",
      "headers": {
        "API-KEY": "<your-dot-mcp-api-key>"
      }
    }
  }
}
```

{% endtab %}

{% tab title="Claude Desktop config file" %}
Edit Claude Desktop's config file:

* **macOS**: `~/Library/Application Support/Claude/claude_desktop_config.json`
* **Windows**: `%APPDATA%\Claude\claude_desktop_config.json`

Paste the Dot server block from the **API Token** tab on Dot's Integrations page, save, and restart Claude Desktop. Look for the MCP indicator (hammer icon) in the chat input.
{% endtab %}
{% endtabs %}

#### Token management

* **Special MCP token with limited scope** — separate from regular API tokens
* **One active token per user** — generating a new token revokes the previous one
* **365-day expiry** — tokens expire after one year
* **Immediate revocation** — admins can delete tokens anytime via the UI

### Asking questions

Once configured, you can ask your AI assistant questions about your data:

* "What were total sales last quarter?"
* "Show me top 10 customers by revenue"
* "Compare this month's performance to last month"
* "What data sources are available?"

### Security

#### OAuth sessions

* OAuth uses industry-standard **OAuth 2.1 with PKCE** for secure authorization
* Access tokens expire after **1 hour** and are automatically refreshed
* Refresh tokens are valid for **1 year** — after that, re-authentication is required
* To end MCP access, disconnect or remove Dot in your client, which revokes that session, or delete your MCP token in Settings. Note that changing your Dot password does not revoke existing MCP connections on its own
* The consent page shows exactly which permissions the client is requesting before you approve

#### Data access

* MCP respects all your Dot permissions and has the same data permissions as the user
* Queries run within your organization's scope and user scope
* All data filtering rules are enforced
* All queries are logged for compliance

### Troubleshooting

#### OAuth issues

**Browser doesn't open for authentication**

* Ensure your AI client is up to date
* Try the token-based method as a fallback
* Check that your browser isn't blocking pop-ups from Dot

**"Authorization session expired" error**

* The consent page expires after 10 minutes — restart the connection from your AI client

**"Sign in to continue" on consent page**

* You need to be signed into Dot in your browser first
* Click "Open Dot Login", sign in (password or SSO), then click "Continue"

#### Token issues

**"Invalid API token" error**

* Verify you copied the complete token
* Check if token was revoked or replaced
* Ensure you're using the correct Dot URL

**"Token does not have required MCP access" error**

* Generate a new MCP token from Settings → Integrations
* Ensure you're not using a regular API token

#### Connection issues

**Connection timeout**

* Check network connection to Dot
* Verify the Dot instance URL is correct
* Some clients have a tool timeout setting — adjust it accordingly


# Databases

Connect Dot to your data warehouse

## Removing a connection

Removing a connection keeps the work you put into your tables. Tables that were switched on are archived instead of deleted, so their descriptions and column comments stay with them.

Reconnect the same warehouse and Dot restores those tables on the next sync, and switches them back on. It matches them by database, schema, and table name, not by the old connection, so re-adding the same warehouse under a different host still works.

Tables that were switched off are removed for good, together with the queries saved against them.

An archived table stays visible on the Model page with a note explaining that it is no longer reachable. From there you can leave it for a future reconnect, or click **Delete Permanently** to remove it.

{% hint style="warning" %}
The remove dialog has a checkbox, **Also delete table setup**. Tick it only if you want the descriptions, the column comments, and the record of which tables were on to be gone for good. Reconnecting does not bring them back.
{% endhint %}


# Snowflake

## Create a Role and a User

This creates a dedicated role and technical user. Replace `example_wh` with your preferred warehouse. `XS` is enough for most installations. It is ok to share this warehouse with other workloads to save costs. **Set a secure password.**

```sql
create role dot_role;
create user dot_user
    password = '<something secret>' -- remember that!
    default_warehouse = example_wh  -- specify your warehouse
    default_role = dot_role;
grant role dot_role to user dot_user;

--allow usage of your warehouse
grant usage on warehouse example_wh to role dot_role;
```

## Grants Read Access to Data

It is recommended to grant permissions only to schemas or tables your end-users should have access to. This is usually a schema with core or reporting tables.

```sql
-- gives access to all objects in a schema 
set db_name = 'example_db'; -- specify name of database 
set schema_name = 'example_schema'; -- specify name of schema 
set db_schema_name = $db_name || '.' || $schema_name; 

grant usage on database identifier($db_name) to role dot_role; 
grant usage on schema identifier($db_schema_name) to role dot_role; 
grant select on all tables in schema identifier($db_schema_name) to role dot_role; 
grant select on future tables in schema identifier($db_schema_name) to role dot_role; 
grant select on all views in schema identifier($db_schema_name) to role dot_role; 
grant select on future views in schema identifier($db_schema_name) to role dot_role; 
grant select on all materialized views in schema identifier($db_schema_name) to role dot_role; 
grant select on future materialized views in schema identifier($db_schema_name) to role dot_role;
```

For shared databases the following statement is enough.

```sql
grant imported privileges on database shared_external_db to role dot_role;
```

## Grants Read Access to Account Information (optional)

Grant access to the query history from Snowflake.

```sql
grant imported privileges on database snowflake to role dot_role;
```

## Allow Dot IPs

If your organization uses a network policy to manage Snowflake access, Dot will only access your Snowflake through the following IPs:

* `5.78.211.110`
* `178.105.217.177`<br>

## Sync Snowflake Roles (optional)

Snowflake can stay the single source of truth for who sees what. The Snowflake connection has two role-sync toggles (both off by default, under **Settings → Connections → Snowflake**):

### Sync Snowflake Roles — tables

On every sync, each table is tagged with the Snowflake roles that can `SELECT` it (from `SHOW GRANTS`), as Dot groups. Only users in a matching group can see and query the table through Dot.

Note: this overwrites the table's existing groups on every sync — Snowflake owns table access from then on.

### Sync Snowflake Roles to Users

The counterpart for people: on every sync, Snowflake users are matched to Dot users by email (the user's `email` or `login_name`), and their granted roles are assigned as Dot groups — using the same group names as the table side, so role-gated tables and role-granted users line up automatically. When someone changes teams in Snowflake, their Dot access follows on the next sync.

* Groups assigned by the sync are tracked separately: a revoked Snowflake role is removed again, while groups you assigned manually in Dot are never touched.
* Disabled Snowflake users are skipped.

Listing users requires extra visibility for the connection role:

```sql
grant manage grants on account to role dot_role;
```

Without it, the user sync skips safely (with a hint in the sync log) and no user is modified — table syncing is unaffected.

## Internal Marketplace (optional)

If your Snowflake organization publishes data products internally, Dot can list them and search them alongside your other tables.

Turn on **Internal Marketplace** under **Settings → Connections → Snowflake**. It is off by default. Turning it on also switches on the Snowflake Marketplace skill under **Settings → Skills**.

Dot reads the organizational listings your connection's role can see, and refreshes them every time the connection syncs. If that refresh fails, the rest of the sync still finishes and the sync log says what went wrong.

Dot only reads the catalog. Mounting a listing's data, and granting your Dot role access to it, still happens in Snowflake.


# BigQuery

## Prerequisites

You’ll need to have a [Google Cloud Platform](https://cloud.google.com/) account with a [project](https://cloud.google.com/storage/docs/projects) you would like Dot to use. Consult the Google Cloud Platform documentation for how to [create and manage a project](https://cloud.google.com/resource-manager/docs/creating-managing-projects). This project should have a BigQuery dataset for Dot to connect to.

## 1 Create a service account

**Create a service account that you manage in your Google Cloud account.** This account should be provisioned with the following read-only roles:

* `bigquery.dataViewer`
* `bigquery.jobUser`
* `bigquery.readSessionUser`

You'll need to provide the service account's email, a [JSON-formatted key](https://cloud.google.com/iam/docs/creating-managing-service-account-keys#creating), and the location of your BigQuery instance.

<details>

<summary>Create a service account step by step.</summary>

1. **Navigate to Service Accounts**:
   * Go to the [Google Cloud Console](https://console.cloud.google.com/).
   * In the **Navigation menu**, select **IAM & Admin >** [**Service Accounts**](https://console.cloud.google.com/iam-admin/serviceaccounts).
2. **Create a New Service Account**:
   * Click on **Create Service Account** at the top.
   * Assign a **Name** and optional **Description** (e.g., `dot-service-account` for identification).
   * Click **Create and Continue**.
3. **Assign Required Roles**:
   * In the **Grant this service account access to project** section, add the following roles:
     * **BigQuery Data Viewer** (`roles/bigquery.dataViewer`)
     * **BigQuery Job User** (`roles/bigquery.jobUser`)
     * **BigQuery Read Session User** (`roles/bigquery.readSessionUser`)
   * Click **Continue** to finalize the role assignments.
4. **Create a JSON Key**:
   * Under **Create key (optional)**, select **JSON** and click **Create**.
   * This downloads a JSON file with the service account credentials. Store this file securely; it contains sensitive information.
5. **Service Account Details Needed for Dot**:
   * **Service Account Email**: Visible in the **Email** column on the Service Accounts page.
   * **JSON Key**: The file downloaded in step 4.
   * **BigQuery Location**: The regional or multi-regional setting for your BigQuery instance (e.g., `us-central1`). Find this in the BigQuery console under **BigQuery > Settings**.

</details>

## 2 Granting permissions

The service account also needs the appropriate read-only roles.

The easiest way to grant these roles is through the [Google Cloud Shell](https://console.cloud.google.com/home/dashboard?cloudshell=true).

First, we'll create a custom role for Dot-related permissions and then bind it to the service account that you're using. We'll also bind read-only BigQuery roles to the service account.

**A) Create a Dot custom role**

```bash
gcloud iam roles create DotMonitor \
  --project={{PROJECT_ID}} \
  --title=DotMonitor \
  --description="Dot specific permissions" \
  --permissions=bigquery.jobs.listAll,bigquery.jobs.list
```

*Note that the* `{{PROJECT_ID}}` *placeholder needs to be replaced with your project id.*

**B) Bind the custom role to a service account and apply read-only BQ roles**

<pre class="language-bash"><code class="lang-bash"><strong>gcloud projects add-iam-policy-binding {{PROJECT_ID}} \
</strong>  --member="serviceAccount:{{SERVICE_ACCOUNT}}" \
  --role="projects/{{PROJECT_ID}}/roles/DotMonitor"

gcloud projects add-iam-policy-binding {{PROJECT_ID}} \
  --member="serviceAccount:{{SERVICE_ACCOUNT}}" \
  --role="roles/bigquery.dataViewer"

gcloud projects add-iam-policy-binding {{PROJECT_ID}} \
  --member="serviceAccount:{{SERVICE_ACCOUNT}}" \
  --role="roles/bigquery.jobUser"

gcloud projects add-iam-policy-binding {{PROJECT_ID}} \
  --member="serviceAccount:{{SERVICE_ACCOUNT}}" \
  --role="roles/bigquery.readSessionUser"
</code></pre>

*Note that the* `{{SERVICE_ACCOUNT}}` *and* `{{PROJECT_ID}}` *placeholders needs to be replaced with your service account and project id, respectively.*

Example Values

* PROJECT\_ID: `super-position-123456`
* SERVICE\_ACCOUNT: `dot-101@super-position-123456.iam.gserviceaccount.com`

## Per-User Access (Optional)

By default, all Dot users in your organization share the same service account when querying BigQuery. If you need each user to only see the data they have access to in BigQuery — based on their individual IAM roles, row-level security policies, or column-level policy tags — you can enable **per-user access** via domain-wide delegation.

When enabled, Dot runs each query as the logged-in user's Google Workspace identity instead of the shared service account. BigQuery enforces access controls natively, so you manage permissions in GCP — not in Dot.

### Prerequisites

* **Google Workspace** — per-user access uses [domain-wide delegation](https://developers.google.com/identity/protocols/oauth2/service-account#delegatingauthority), which requires a Google Workspace domain.
* **Recommended: Google SSO** configured in Dot (see [Google SSO setup](/integrations/sso/google)). SSO guarantees that the user's Dot email matches their Google Workspace identity.
* Users who sign in with a password are not delegated by default. If your password-login users have Dot emails that match their Google Workspace emails, you can enable the **Include non-SSO users** sub-toggle.

### Step 1: Enable domain-wide delegation for the service account

1. Go to the [Google Cloud Console](https://console.cloud.google.com/) > **IAM & Admin** > **Service Accounts**.
2. Click on the Dot service account.
3. Under **Show domain-wide delegation**, check **Enable Google Workspace Domain-wide Delegation**.
4. Note the **Client ID** shown (you'll need it in the next step).

### Step 2: Authorize the service account in Google Workspace

1. Go to [Google Workspace Admin Console](https://admin.google.com/) > **Security** > **API Controls** > **Manage Domain Wide Delegation**.
2. Click **Add new**.
3. Enter the **Client ID** from Step 1.
4. Enter the following OAuth scope: `https://www.googleapis.com/auth/bigquery`
5. Click **Authorize**.

### Step 3: Grant BigQuery access to your users

Each user who will use Dot needs BigQuery permissions on the relevant projects and datasets. At minimum:

```bash
gcloud projects add-iam-policy-binding {{PROJECT_ID}} \
  --member="user:alice@yourcompany.com" \
  --role="roles/bigquery.dataViewer"

gcloud projects add-iam-policy-binding {{PROJECT_ID}} \
  --member="user:alice@yourcompany.com" \
  --role="roles/bigquery.jobUser"
```

For managing permissions at scale, use [Google Groups](https://cloud.google.com/iam/docs/groups-in-cloud-console) — grant roles to a group and add users to it.

{% hint style="info" %}
You can also use BigQuery [row-level security policies](https://cloud.google.com/bigquery/docs/row-level-security-intro) and [column-level security (policy tags)](https://cloud.google.com/bigquery/docs/column-level-security-intro) for fine-grained access control. These are enforced automatically when per-user access is enabled.
{% endhint %}

### Step 3: Enable per-user access in Dot

1. Go to **Settings** > **Connections** > **BigQuery**.
2. Click **Edit**.
3. Enable the **Per-user BigQuery access** toggle.
4. Optionally enable **Include non-SSO users** if your password-login users should also be impersonated.
5. Click **Connect** to save.

Once enabled, every query a user runs in Dot will execute as their Google identity. If a user doesn't have access to a table or column in BigQuery, they'll see a clear error message instead of the data.

### How it works

| Scenario                       | Who runs the query                                                                      |
| ------------------------------ | --------------------------------------------------------------------------------------- |
| User logged in via Google SSO  | The user's Google identity                                                              |
| User logged in with password   | The shared service account (or user's identity if **Include non-SSO users** is enabled) |
| Scheduled queries and alerts   | The schedule creator's Google identity                                                  |
| Data sync and model operations | The shared service account                                                              |

{% hint style="warning" %}
**Scheduled queries run as the creator.** If a user's Google account is deactivated (e.g., they leave the company), their scheduled queries will fail. After 3 consecutive failures, the schedule is automatically paused and the owner is notified via email. To fix this, reassign the schedule to an active user.
{% endhint %}

### Known limitations

* **Table metadata is shared.** Data sync (schema discovery, AI-generated descriptions) always runs as the shared service account. This means all users can see table names, column names, and AI-generated descriptions for all synced tables — even tables they cannot query. No actual data values are exposed, but the table structure and descriptions are visible to all users in the organization.
* **BigQuery row-level security returns empty results.** If a user has table access but is restricted by BigQuery row-level security policies, queries return empty results rather than an error. Dot will inform the user that no data was found but cannot distinguish between "no matching rows" and "access restricted."

## Allow Dot IPs

If your organization uses a network policy to manage BigQuery access, Dot will only access your BigQuery through the following IPs:

* `5.78.211.110`
* `178.105.217.177`


# Redshift

## Create a Group and a User

As a super user, execute the following SQL commands to create a group, a user assigned to that group, and permissions to access a system table.

```sql
CREATE USER dot_user PASSWORD '<something secret>' SYSLOG ACCESS UNRESTRICTED;
ALTER USER dot_user SYSLOG ACCESS UNRESTRICTED;

CREATE GROUP dot_group;

ALTER GROUP dot_group ADD USER dot_user;

-- Grant select to system table for meta data
GRANT SELECT ON svv_table_info TO GROUP dot_group;
```

## Grants Read Access to Data

Then for each schema `schema`, execute the following three commands to grant read-only access.

```sql
-- Grant usage on schema and select on current and future child tables
GRANT USAGE ON SCHEMA "schema" TO GROUP dot_group;
GRANT SELECT ON ALL TABLES IN SCHEMA "schema" TO GROUP dot_group;
ALTER DEFAULT PRIVILEGES IN SCHEMA "schema" GRANT SELECT ON TABLES TO GROUP dot_group;
```

**Note**: to programmatically generate these three queries for all schemas, you can use the following command. These commands still need to be executed.

```sql
SELECT 
    'GRANT USAGE ON SCHEMA "' || schema_name || '" TO GROUP dot_group;' || '\n' ||
    'GRANT SELECT ON ALL TABLES IN SCHEMA "' || schema_name || '" TO GROUP dot_group;' || '\n' ||
    'ALTER DEFAULT PRIVILEGES IN SCHEMA "' || schema_name || '" GRANT SELECT ON TABLES TO GROUP dot_group;' AS single_schema_statement
FROM svv_all_schemas
WHERE schema_name not in ('information_schema', 'pg_catalog', 'pg_internal');
```

## Allow Dot IPs

If your organization uses a network policy to manage Redshift access, Dot will only access your Redshift through the following IPs:

* `5.78.211.110`
* `178.105.217.177`

If you'd rather not make the cluster publicly accessible, you can skip the steps below and use an [SSH tunnel](#connect-via-ssh-tunnel) instead.

1. In the Redshift dashboard, **click on the desired cluster name**.

<div align="left"><figure><img src="https://files.readme.io/39a5a42-1.png" alt="" width="375"><figcaption></figcaption></figure></div>

2. When viewing information for your Redshift cluster, click the **Properties** tab.

![1918](https://files.readme.io/fcb267a-2.png)

3. Scroll down to the **Network and security settings** section.

![1918](https://files.readme.io/7ed17e1-Screen_Shot_2021-04-19_at_3.13.51_PM.png)

4. If **Public Accessibility** is not enabled, click **Edit publicly accessible** button then enable.

![1914](https://files.readme.io/a7472c2-4.png) ![1918](https://files.readme.io/a18855d-3_-_publicly_accessible.png)

5. Click **VPC security group link**.

![1918](https://files.readme.io/c321c34-3.png)

6. Click **Edit inbound rules**.

![1918](https://files.readme.io/fcced0a-5.png)

7. Add the following IPs of Type `Redshift`:

* `5.78.211.110`
* `178.105.217.177`

![](/files/VJusHJJtEebPSYw3ejUE)<br>

## Connect via SSH Tunnel

If you'd rather not make the cluster publicly accessible, Dot can reach a private Redshift cluster through an SSH bastion host in your VPC instead. In the Redshift connection dialog, enable **Connect via SSH Tunnel** and provide:

* **SSH Host**: your bastion / jump server
* **SSH Port**: `22` (default) or a custom port
* **SSH Username**: the SSH user
* **SSH Authentication**: SSH password or private key

Dot tunnels the Redshift connection through the bastion, so the cluster keeps its private endpoint and never needs public accessibility.


# AWS Athena

To integrate Amazon Athena with Dot securely and efficiently, follow these steps to establish a connection, configure IAM permissions, and restrict access to specific Athena workgroups.

**1. Create a Dedicated IAM User for Dot**

For enhanced security, it's advisable to create a dedicated IAM user for Dot with permissions limited to the necessary Athena resources.

* **Create the IAM User**:
  * Sign in to the AWS Management Console.
  * Navigate to the IAM service.
  * Select "Users" and then "Add user".
  * Enter a username (e.g., `dot_athena_user`).
  * Choose "Programmatic access" to provide access via the AWS CLI, SDKs, or APIs.
* **Attach Policies**:
  * Select "Attach existing policies directly".
  * Attach the `AmazonAthenaFullAccess` managed policy to grant full access to Athena.
  * Attach the `AWSGlueConsoleFullAccess` policy to allow access to the AWS Glue Data Catalog, which Athena uses for metadata.
  * Ensure the user has the necessary permissions to access the relevant Amazon S3 buckets where your data resides.
* **Complete User Creation**:
  * Review the permissions and create the user.

**2. Optinonally restrict IAM Permissions to Specific Athena Workgroups**

To enhance security, you can limit the IAM user's permissions to specific Athena workgroups. This ensures Dot accesses only the intended resources.

* **Define the Policy**:
  * Create a custom IAM policy that grants permissions solely to the desired workgroups.
  * Specify the workgroup ARNs in the policy's `Resource` section.
* **Example Policy**:

```json
  {
    "Version": "2012-10-17",
    "Statement": [
      {
        "Effect": "Allow",
        "Action": [
          "athena:StartQueryExecution",
          "athena:StopQueryExecution",
          "athena:GetQueryExecution",
          "athena:GetQueryResults"
        ],
        "Resource": [
          "arn:aws:athena:us-east-1:123456789012:workgroup/your_workgroup_name"
        ]
      },
      {
        "Effect": "Allow",
        "Action": [
          "glue:GetDatabase",
          "glue:GetTable",
          "glue:GetPartitions"
        ],
        "Resource": "*"
      },
      {
        "Effect": "Allow",
        "Action": [
          "s3:GetObject",
          "s3:ListBucket"
        ],
        "Resource": [
          "arn:aws:s3:::your_bucket_name",
          "arn:aws:s3:::your_bucket_name/*"
        ]
      }
    ]
  }

```

 Replace `your_workgroup_name` with the name of your Athena workgroup and `your_bucket_name` with your S3 bucket name.

* **Attach the Policy**:
  * Navigate to the IAM console.
  * Select "Policies" and then "Create policy".
  * Use the JSON editor to input your custom policy.
  * Review and create the policy.
  * Attach this policy to the IAM user created for Dot.

For more details on controlling workgroup access with IAM policies, refer to the [AWS Athena documentation](https://docs.aws.amazon.com/athena/latest/ug/workgroups-iam-policy.html).

**3. Configure Network Access**

Ensure that Dot can communicate with Athena by allowing outbound access to the following IP addresses, especially if your organization uses a firewall or network policies:

* `5.78.211.110`
* `178.105.217.177`

**4. Obtain the Access Key ID and Secret Access Key**

After creating the IAM user, generate the access keys required for programmatic access:

* **Generate Access Keys via AWS Management Console**:
  * In the IAM console, select "Users" and click on the username you created (e.g., `dot_athena_user`).
  * Navigate to the "Security credentials" tab.
  * In the "Access keys" section, click "Create access key".
  * Choose the appropriate use case (e.g., "Command Line Interface (CLI)") and click "Next".
  * Optionally, add a description tag for the access key, then click "Create access key".
  * The Access Key ID and Secret Access Key will be displayed. **This is the only time the Secret Access Key will be available**, so ensure you securely store it, either by downloading the `.csv` file or copying the keys to a secure location.

**5. Connect Dot to Amazon Athena**

With the IAM user and network configurations in place, proceed to connect Dot to Athena:

* Navigate to Dot's integration settings.
* Select "Add Integration" and choose Amazon Athena from the list of available data sources.
* Enter the Access Key ID and Secret Access Key of the IAM user created earlier.
* Specify the AWS region where your Athena instance is located.
* Provide any additional configuration details as prompted by Dot.

By following these steps, Dot will establish a secure connection to Amazon Athena, enabling efficient data access while adhering to security best practices.


# Databricks

## Connect Dot to Databricks

### 1. Generate a Databricks Access Token

Dot requires an access token to connect to Databricks. It's recommended to use a service principal for this purpose.

#### Option A: Using a Service Principal

1. **Create a Service Principal**\
   Follow the [Databricks documentation](https://docs.databricks.com/api/workspace/serviceprincipals/create) to create a service principal.
2. **Grant Token Usage to the Service Principal**\
   Ensure the service principal has permissions to use access tokens.
3. **Generate an Access Token**\
   Use the Databricks CLI to generate an access token:

   ```bash
   databricks tokens create --comment "Dot Integration" --lifetime-seconds 0
   ```

Save the generated token securely.

#### Option B: Using a Personal Access Token

Alternatively, you can generate a personal access token for your user account.

### 2. Grant Data Permissions to Dot's Service Principal

To allow Dot to access the necessary data, grant the appropriate permissions.

#### Unity Catalog

**Access to All Tables in a Catalog**

```sql
GRANT USE CATALOG ON CATALOG <catalog_name> TO `<application_id>`;
GRANT USE SCHEMA ON CATALOG <catalog_name> TO `<application_id>`;
GRANT SELECT ON CATALOG <catalog_name> TO `<application_id>`;
```

**Access to Specific Tables**

```sql
GRANT USE CATALOG ON CATALOG <catalog_name> TO `<application_id>`;
GRANT USE SCHEMA ON SCHEMA <catalog_name>.<schema_name> TO `<application_id>`;
GRANT SELECT ON TABLE <catalog_name>.<schema_name>.<table_name> TO `<application_id>`;
```

**Access to System Tables for Data Insights**

1. **Enable System Schemas**\
   Use the Databricks API to enable system schemas:

   ```bash
   curl -X PUT -H "Authorization: Bearer <token>" \
   "https://<workspace_url>/api/2.0/unity-catalog/metastores/<metastore_id>/systemschemas/query"
   ```
2. **Grant Access to System Tables**

   ```sql
   GRANT USE SCHEMA ON SCHEMA system.query TO `<application_id>`;
   GRANT SELECT ON TABLE system.query.history TO `<application_id>`;
   GRANT USE SCHEMA ON SCHEMA system.access TO `<application_id>`;
   GRANT SELECT ON TABLE system.access.column_lineage TO `<application_id>`;
   ```

#### Hive Metastore

```sql
GRANT READ_METADATA, USAGE, SELECT ON CATALOG <catalog_name> TO `<application_id>`;
```

### 3. Create a Databricks SQL Warehouse for Dot

1. **Create a SQL Warehouse**\
   Follow the [Databricks documentation](https://docs.databricks.com/compute/sql-warehouse/create.html) to create a SQL Warehouse.
2. **Assign Permissions**\
   In the SQL Warehouse settings, click 'Permissions' and grant the Dot service principal 'Can Use' permissions.

### 4. Add Databricks as a Connection in Dot

1. **Navigate to Connections**\
   In Dot, go to the Settings / Connections page.
2. **Add a New Connection**\
   Click '+ Database Connection' and select Databricks
3. **Enter Connection Details**\
   Provide the following information:
   * **Host**: From the SQL Warehouse created in Step 3
   * **Port**: Typically 443
   * **HTTP Path**: From the SQL Warehouse
   * **Access Token**: Generated in Step 1

***

This guide should help you set up a connection between Dot and Databricks.


# Postgres

## Create a Group and a User

As a super user, execute the following SQL commands to create a group, a user assigned to that group, and permissions to access a system table.

```sql
CREATE USER dot_user PASSWORD '<something secret>';

CREATE GROUP dot_group;

ALTER GROUP dot_group ADD USER dot_user;

-- Grant Postgres' monitor role for meta data
GRANT pg_monitor TO dot_group
```

## Grants Read Access to Data

Then for each schema `schema`, execute the following three commands to grant read-only access.

```sql
-- Grant usage on schema and select on current and future child tables
GRANT USAGE ON SCHEMA "schema" TO GROUP dot_group;
GRANT SELECT ON ALL TABLES IN SCHEMA "schema" TO GROUP dot_group;
ALTER DEFAULT PRIVILEGES IN SCHEMA "schema" GRANT SELECT ON TABLES TO GROUP dot_group;
```

**Note**: to programmatically generate these three queries for all schemas, you can use the following command. These commands still need to be executed.

```sql
SELECT 
    'GRANT USAGE ON SCHEMA "' || schema_name || '" TO GROUP dot_group;' || '\n' ||
    'GRANT SELECT ON ALL TABLES IN SCHEMA "' || schema_name || '" TO GROUP dot_group;' || '\n' ||
    'ALTER DEFAULT PRIVILEGES IN SCHEMA "' || schema_name || '" GRANT SELECT ON TABLES TO GROUP dot_group;' AS single_schema_statement
FROM svv_all_schemas
WHERE schema_name not in ('information_schema', 'pg_catalog', 'pg_internal');
```

## Allow Dot IPs

If your organization uses a firewall to manage Postgres access, Dot will only access your Postgres through the following IPs:

* `5.78.211.110`
* `178.105.217.177`

## Connect via SSH Tunnel

If your database is in a private network or behind a firewall, Dot can connect through an SSH bastion host instead of exposing it directly. In the connection dialog, enable **Connect via SSH Tunnel** and provide:

* **SSH Host**: your bastion / jump server
* **SSH Port**: `22` (default) or a custom port
* **SSH Username**: the SSH user
* **SSH Authentication**: SSH password or private key

Dot tunnels the database connection through the bastion, so the database itself never needs a public endpoint.


# Microsoft SQL Server

## **Create a Login and a User**

To set up access in SQL Server, one would create a login at the server level and then a user at the database level linked to that login. Below is an equivalent SQL Server script:

```sql
-- Replace placeholder values (YourDatabaseName, YourSchemaName, and <something secret>) 
-- with your actual values.

-- Create server-level login
CREATE LOGIN dot_login WITH PASSWORD = '<something secret>';

-- Use the desired database
USE YourDatabaseName;

-- Create database-level user linked to the server login
CREATE USER dot_user FOR LOGIN dot_login;

-- Add user to a role (for demonstration purposes, using db_datareader which gives read-only access)
ALTER ROLE db_datareader ADD MEMBER dot_user;

```

## **Grants Read Access to Data**

In SQL Server, the **`db_datareader`** role grants read-only access to all current and future tables in the database. If you've added the user to this role, you don't need to do schema-specific grants. However, if you wish to grant read access only on specific schemas:

```sql
-- Grant select on all current tables in a schema
GRANT SELECT ON SCHEMA::YourSchemaName TO dot_user;

-- Note: SQL Server does not support the concept of default privileges.
-- Any new tables or views would need permissions assigned explicitly.

```

## **Allow Dot IPs**

If you are using SQL Server's firewall features or another firewall mechanism to control access, ensure that you whitelist the following IPs to allow Dot to access your SQL Server:

* 5.78.211.110
* 178.105.217.177


# Microsoft Fabric

Connect Dot to your Microsoft Fabric Data Warehouse or SQL Analytics Endpoint to query your data using natural language.

## Authentication Methods

Microsoft Fabric supports two authentication methods for external applications like Dot:

### Option 1: Service Principal Authentication (Recommended)

Service Principal authentication provides secure, automated access without requiring user credentials.

**1. Create an Entra ID Application and Service Principal**

1. **Register a new application** in Azure Portal:
   * Go to Azure Active Directory > App registrations > New registration
   * Name: `Dot-Fabric-Access` (or similar)
   * Account types: Accounts in this organizational directory only
   * Click "Register"
2. **Create a client secret**:
   * Go to Certificates & secrets > New client secret
   * Description: `Dot Integration`
   * Expires: Choose appropriate duration (12 months recommended)
   * Copy the **Value** (this is your `client_secret`)
3. **Note the following values** from the Overview page:
   * **Application (client) ID** (this is your `client_id`)
   * **Directory (tenant) ID** (this is your `tenant_id`)

**2. Configure Fabric Tenant Settings**

1. **Enable Service Principal API access**:
   * Go to Fabric Admin Portal > Tenant settings
   * Find "Developer settings" > "Service principals can use Fabric APIs"
   * Enable this setting for your organization or specific security groups

**3. Grant Workspace Access**

1. **Add Service Principal to Workspace**:
   * Go to your Fabric workspace
   * Settings > Manage access
   * Add people or groups > Enter your Service Principal name
   * Assign **Contributor** role (recommended) or **Member** role

### Option 2: Entra UPN Authentication

Use your Microsoft Entra ID username and password for direct authentication.

**Requirements:**

* Valid Entra ID user account with access to the Fabric workspace
* Username in UPN format (e.g., `user@company.com`)
* Account password

**Grant User Access:**

1. **Add user to Workspace**:
   * Go to your Fabric workspace
   * Settings > Manage access
   * Add the user account
   * Assign **Contributor** role (recommended)

## Connection Information

### Required Connection Details

**For Service Principal Authentication:**

* **Server**: Your Fabric SQL Analytics Endpoint (e.g., `abc123def456.datawarehouse.fabric.microsoft.com`)
* **Database**: Your warehouse database name
* **Tenant ID**: Directory (tenant) ID from Azure AD
* **Client ID**: Application (client) ID from Azure AD
* **Client Secret**: Secret value created in Azure AD

**For Entra UPN Authentication:**

* **Server**: Your Fabric SQL Analytics Endpoint
* **Database**: Your warehouse database name
* **Username**: Your Entra UPN (e.g., `user@company.com`)
* **Password**: Your account password

#### Finding Your Connection String

1. **Get your SQL Analytics Endpoint**:
   * Open your Fabric workspace
   * Go to your Data Warehouse or Lakehouse
   * Settings > SQL endpoint
   * Copy the server name from the connection string

### Network Requirements

#### Dot IP Allowlist

If your organization uses network firewalls or IP restrictions, add these Dot service IP addresses to your allowlist:

* `5.78.211.110`
* `178.105.217.177`

#### Port Requirements

* **Port 1433** (TCP) - Standard SQL Server port for TDS connections
* **Port 443** (TCP) - HTTPS for Fabric API access


# Clickhouse

Connect Dot to your ClickHouse database to query your data using natural language.

## Create a User for Dot

#### 1. Create a Dedicated User

Create a dedicated user for Dot with read-only access to your data:

```sql
-- Create a user for Dot
CREATE USER dot_user IDENTIFIED BY 'your_secure_password_here';

-- Grant read access to specific databases
GRANT SELECT ON database_name.* TO dot_user;

-- Or grant read access to all databases (if needed)
-- GRANT SELECT ON *.* TO dot_user;
```

#### 2. Grant Access to System Tables (Optional)

For enhanced query analytics and usage insights, grant access to system tables:

```sql
-- Grant access to query history for usage analytics
GRANT SELECT ON system.query_log TO dot_user;
GRANT SELECT ON system.tables TO dot_user;
GRANT SELECT ON system.columns TO dot_user;
GRANT SELECT ON system.databases TO dot_user;
```

This enables Dot to:

* Analyze table usage patterns
* Identify popular queries
* Provide performance optimization suggestions
* Show data lineage insights

## Connection Methods

### Option 1: Direct HTTP/HTTPS Connection (Recommended)

For most ClickHouse deployments, use the HTTP interface with SSL encryption:

**Required Connection Details:**

* **Host**: Your ClickHouse server hostname (e.g., `clickhouse.company.com`)
* **Port**: `8443` (HTTPS) or `8123` (HTTP - not recommended for production)
* **Username**: The user you created (e.g., `dot_user`)
* **Password**: The secure password you set
* **Database**: Your target database name (optional - can access all granted databases)
* **SSL**: Enable for production (`secure=true`)
* **Certificate Verification**: Enable unless using self-signed certificates

**Connection Example:**

```
Host: clickhouse.company.com
Port: 8443
Username: dot_user
Password: your_secure_password
Database: analytics_db
```

### Option 2: SSH Tunnel Connection

For databases behind firewalls or in private networks, use SSH tunneling:

**Required SSH Details:**

* **SSH Host**: Your bastion/jump server
* **SSH Port**: `22` (default) or custom SSH port
* **SSH Username**: Your SSH username
* **SSH Authentication**: Password or private key

**SSH Configuration Example:**

```
SSH Host: bastion.company.com
SSH Port: 22
SSH Username: your_ssh_user
SSH Authentication: Private Key
ClickHouse Host: internal-clickhouse.local
ClickHouse Port: 8443
```

## Dot IP Allowlist

If your ClickHouse server uses IP restrictions, add these Dot service IP addresses to your allowlist:

* `5.78.211.110`
* `178.105.217.177`

#### Port Requirements

**Standard Ports:**

* **Port 8443** (TCP) - HTTP interface with SSL/TLS (recommended)
* **Port 8123** (TCP) - HTTP interface without encryption (not recommended)
* **Port 9440** (TCP) - Native protocol with SSL/TLS
* **Port 9000** (TCP) - Native protocol without encryption

**SSH Tunnel Ports:**

* **Port 22** (TCP) - SSH for tunneling (if using SSH option)


# MySQL / MariaDB

### Create a User and Assign Privileges

As a superuser, execute the following SQL commands to create a user and assign necessary privileges.

```sql
CREATE USER 'dot_user'@'%' IDENTIFIED BY '<something secret>';

CREATE ROLE 'dot_role';

GRANT 'dot_role' TO 'dot_user'@'%';
```

### Grant Read Access to Data

For each database (`database`), execute the following commands to grant read-only access.

```sql
GRANT USAGE ON `database`.* TO 'dot_role';
GRANT SELECT ON `database`.* TO 'dot_role';
```

**Note**: To programmatically generate these grant statements for all databases, you can use the following command. These commands still need to be executed.

```sql
SELECT 
    CONCAT(
        'GRANT USAGE ON `', schema_name, '`.* TO \'dot_role\';',
        '\nGRANT SELECT ON `', schema_name, '`.* TO \'dot_role\';'
    ) AS grant_statements
FROM information_schema.schemata
WHERE schema_name NOT IN ('mysql', 'information_schema', 'performance_schema', 'sys');
```

### Allow Dot IPs

If your organization uses a firewall to manage MySQL access, Dot will only access your MySQL through the following IPs:

* `5.78.211.110`
* `178.105.217.177`

### Connect via SSH Tunnel

If your database is in a private network or behind a firewall, Dot can connect through an SSH bastion host instead of exposing it directly. In the connection dialog, enable **Connect via SSH Tunnel** and provide:

* **SSH Host**: your bastion / jump server
* **SSH Port**: `22` (default) or a custom port
* **SSH Username**: the SSH user
* **SSH Authentication**: SSH password or private key

Dot tunnels the database connection through the bastion, so the database itself never needs a public endpoint.


# Oracle Database

### Create a User

As a DBA or SYSDBA, create a dedicated read-only user for Dot.

```sql
CREATE USER dot_reader IDENTIFIED BY '<something secret>';
GRANT CREATE SESSION TO dot_reader;
```

### Grant Read Access to Data

For each schema you want Dot to access, grant SELECT on all tables.

```sql
BEGIN
  FOR t IN (SELECT table_name FROM all_tables WHERE owner = '<SCHEMA>') LOOP
    EXECUTE IMMEDIATE 'GRANT SELECT ON <SCHEMA>.' || t.table_name || ' TO dot_reader';
  END LOOP;
END;
/
```

**Note**: Replace `<SCHEMA>` with the schema/owner name (e.g. `HR`, `SALES`). Repeat for each schema you want to expose to Dot.

To also grant access to views:

```sql
BEGIN
  FOR v IN (SELECT view_name FROM all_views WHERE owner = '<SCHEMA>') LOOP
    EXECUTE IMMEDIATE 'GRANT SELECT ON <SCHEMA>.' || v.view_name || ' TO dot_reader';
  END LOOP;
END;
/
```

### Connection Details

In Dot, use these fields:

* **Host**: Your Oracle server hostname or IP address
* **Port**: `1521` (default Oracle listener port)
* **Service Name**: Your Oracle service name (e.g. `ORCL`, `XEPDB1`, `FREEPDB1`)
* **Username**: `dot_reader`
* **Password**: The password you set above

Dot connects directly using Oracle's TNS protocol (Thin mode) — no Oracle Client installation is required on the Dot side.

### Allow Dot IPs

If your organization uses a firewall to manage Oracle access, Dot will only access your database through the following IPs:

* `5.78.211.110`
* `178.105.217.177`

### Connect via SSH Tunnel

If your database is in a private network or behind a firewall, Dot can connect through an SSH bastion host instead of exposing it directly. In the connection dialog, enable **Connect via SSH Tunnel** and provide:

* **SSH Host**: your bastion / jump server
* **SSH Port**: `22` (default) or a custom port
* **SSH Username**: the SSH user
* **SSH Authentication**: SSH password or private key

Dot tunnels the database connection through the bastion, so the database itself never needs a public endpoint.


# Motherduck & DuckDB

## Motherduck

1. Get your token on [Motherduck Settings](https://app.motherduck.com/settings)

<figure><img src="/files/JIBCei4qYxVdsXHlzs6Q" alt=""><figcaption></figcaption></figure>

2. Compose your connection string

   It follows this pattern. `md:<database_name>?motherduck_token=<your_token>`\
   e.g.: `md:sample_data?motherduck_token=eyABC...`

**Notes**

The Motherduck connector requires DuckDB to be at least v1.4.0.

## DuckDB

If you have your data in a local duckdb file you can either host it yourself (e.g. on S3) or [talk to us](https://github.com/Snowboard-Software/snowboard_software/tree/main/dot/dot/integrations/support.md).


# SAP HANA

Easily talk to your "Systems, Applications, and Products" In-Memory database

### Create a role and a user for Dot <a href="#create-a-role-and-a-user-for-dot" id="create-a-role-and-a-user-for-dot"></a>

Create a dedicated user and role for Dot to access SAP HANA Cloud tables. Replace `<schema_name>`, `<table_name>` with the name of your schema and tables.

```sql
SET SCHEMA <schema_name>
CREATE USER DOT_USER
CREATE ROLE DOT_ROLE
GRANT SELECT ON SCHEMA <schema_name> TO DOT_ROLE
GRANT DOT_ROLE TO DOT_USER
```

It is recommended to grant permissions only to tables your end-users should have access to. These typically include core business tables and reporting tables.

If you need to grant access to a particular table, run the following command:

```sql
GRANT SELECT ON <schema_name>.<table_name> TO <user_or_role>
```

### How to connect to the free SAP HANA Cloud instance <a href="#how-to-connect-to-the-free-sap-hana-cloud-instance" id="how-to-connect-to-the-free-sap-hana-cloud-instance"></a>

1. Sign-in to the SAP BTP Cockpit with your credentials
2. Select your subaccount under the account explorer
3. Select your space under Cloud Foundry -> Spaces
4. Select **SAP HANA Cloud** on the left panel
5. Open **SAP HANA Cloud Central**, check the current status of the SAP HANA instance.

For more info on how to start with the SAP HANA free instance check out: [Provision an Instance of SAP HANA Cloud, SAP HANA Database](https://developers.sap.com/tutorials/hana-cloud-mission-trial-3.html).

### Allow Dot IPs: <a href="#allow-dot-ips" id="allow-dot-ips"></a>

To connect your SAP HANA Cloud instance with Dot, configure it to allow Dot to access it through the following IPs.

* 5.78.211.110
* 178.105.217.177


# Google Sheets

Paste a link and query your spreadsheet like a database table — always the live values.

A lot of business truth never makes it into the warehouse: budgets, targets, headcount plans, campaign trackers, hand-maintained mappings. Connect those sheets and Dot can answer with them — instead of pretending they don't exist.

There is no export-import loop. Dot reads the sheet live on every query, so answers always reflect what the spreadsheet says right now.

## Prerequisites

* A Google Sheet shared as **Anyone with the link → Viewer**. Dot reads sheets through Google's CSV export, so there is no service account to create and no OAuth flow.
* Data laid out as a table: headers in the first row, one record per row.

## Connect in Dot

1. In Google Sheets, click **Share → Anyone with the link → Viewer** and copy the link.
2. In Dot, go to **Settings → Connections** and choose **Google Sheets**.
3. Paste one or more spreadsheet URLs and connect.

Dot discovers every tab in each spreadsheet, skips the ones that aren't shaped like a table, and creates one table per remaining tab — with a short, descriptive name inferred from the content. Names of tabs you've already synced stay stable across re-connects, so notes and saved questions keep working.

{% hint style="info" %}
Odd delimiters, European decimal commas, extra header rows — Dot samples each tab and configures parsing automatically. If a tab comes out wrong, fix the layout in the sheet and sync again.
{% endhint %}

{% hint style="warning" %}
**Anyone with the link means exactly that.** Anyone holding the URL can view the sheet. Use this for data you'd be comfortable circulating internally, and keep sensitive data in a governed source instead.
{% endhint %}

## When to sync

Editing rows and cells needs no sync, because every query reads the sheet live.

Changing the sheet's *shape* does. Adding, renaming or removing a column changes the table Dot queries, and a new tab is a new table, so tell Dot about it:

* Click **Sync** on the connection. Modelers use **Sync now** in the connection panel.
* Or turn on **Schedule sync** in the connection panel to refresh it daily or weekly, the same as a warehouse.

If a tab stops being shared with "Anyone with the link", the sync log says how many sheets could not be refreshed. Re-share it and sync again.

## Good to know

* **Always fresh** — queries read the live sheet; there is no cached copy to go stale.
* **A table like any other** — sheets appear in your Model where you can activate them, describe them, and use them in questions and dashboards alongside your other sources.


# Firebolt

Connect Dot to Firebolt and ask questions over your sub-second analytics engine.

Firebolt is built for interactive, sub-second analytics. Pairing it with Dot means that speed isn't reserved for the people who write SQL — anyone can ask in plain language, and the engine's performance carries straight through to their answer.

## Create a service account for Dot

Dot authenticates with a Firebolt service account:

1. In Firebolt, go to **Configure → Service accounts** and create one for Dot.
2. Attach it to a user whose role has read access to the data you want to expose (`SELECT` on the relevant database plus `USAGE` on the engine — a read-only role is all Dot needs).
3. Note the **client ID** and **client secret**, plus your **account name**, the **database**, and the **engine** Dot should use.

## Connect in Dot

1. Go to **Settings → Connections** and choose **Firebolt**.
2. Enter the client ID, client secret, account name, database, and engine.
3. Connect. Dot syncs your tables and columns, generates descriptions, and is ready for questions.

{% hint style="info" %}
Dot only issues `SELECT` queries — a read-only role is the right scope. Table and column selection, descriptions, and everything else works the same as for any other database in your Model.
{% endhint %}


# Semantic Layers

Connect Dot to your semantic layer so answers use your own metric definitions

If your team already defines metrics in a semantic layer, connect it to Dot. Instead of guessing what a metric means, Dot uses your definitions, so the numbers it gives back match the ones your semantic layer produces.

Pick the one you use:

* [dbt Semantic Layer](/integrations/semantic-layers/dbt-semantic-layer)
* [PowerBI Semantic Layer](/integrations/semantic-layers/powerbi-semantic-layer)
* [Azure Analysis Services](/integrations/semantic-layers/azure-analysis-services)
* [Looker](/integrations/semantic-layers/looker)
* [Malloy](/integrations/semantic-layers/malloy)
* [Steep](/integrations/semantic-layers/steep)


# dbt Semantic Layer

Tell Dot about your most important metrics.

{% embed url="<https://www.loom.com/share/4849e276a2f649a5b84538b062a33320?sid=dd7fa11e-a404-4274-a097-3045b65003f2>" fullWidth="true" %}
Demo of Dot on dbt
{% endembed %}

To connect you need admin access to **dbt cloud** and you need to have [setup the semantic layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/quickstart-sl).

## Create API token

1. Go to ⚙️ | Accout Settings

<figure><img src="/files/wXzobueO7gWpRw5cLgHN" alt=""><figcaption></figcaption></figure>

2. Click on Service Tokens

<div align="center"><figure><img src="/files/dqcWFif27OO8JTldn7qT" alt="" width="548"><figcaption></figcaption></figure></div>

3. Create new Token with permissions for `Semantic Layer` and `Metadata`

<figure><img src="/files/CZj09uCWSX8bOMD4Y9xh" alt=""><figcaption></figcaption></figure>

4. Copy and Save Generated Token

<figure><img src="/files/gdFCIdcgouu4gXgbzMJX" alt=""><figcaption></figcaption></figure>

## Get Environment ID

1. Go to Environments

<figure><img src="/files/LLooccvw0I1mITM4NvQl" alt=""><figcaption></figcaption></figure>

2. Click on your production environment and copy the last part of the url ( e.g. 12345)

<figure><img src="/files/aV8QNUBZ6C1HOb6SJIzL" alt=""><figcaption></figcaption></figure>

## Get Your GraphQL URL

\
Depending on where your dbt account is hosted, you need to obtain a different url.

Here are the official docs on the different schema explorer/Graph QL URLs

{% embed url="<https://docs.getdbt.com/docs/dbt-cloud-apis/sl-graphql#dbt-semantic-layer-graphql-api>" %}

Example: `https://semantic-layer.cloud.getdbt.com/api/graphql`

## Allow Dot IPs

If your organization uses a firewall to manage dbt access, Dot will only access your dbtthrough the following IPs:

* `5.78.211.110`
* `178.105.217.177`


# PowerBI Semantic Layer

Connect Dot to Power BI semantic models with one click

Connect Dot to your Power BI semantic models to query your measures and dimensions using natural language. Dot syncs your workspaces, datasets, measures, and columns—including their descriptions—so you can ask questions about your data without writing DAX.

{% hint style="info" %}
**Requirements**

* **Power BI Premium or Premium Per User (PPU)** capacity is required
* "Dataset Execute Queries REST API" must be enabled in your tenant settings
* You must have access to the workspaces you want to sync
  {% endhint %}

## Connect to Power BI

<figure><img src="/files/Eqjjvg59xxLH1PrPk3eW" alt="Power BI connection settings"><figcaption></figcaption></figure>

1. Go to **Settings → Semantic Layers → Power BI**
2. (Optional) Enter a **Workspace ID** to sync only that workspace, or leave it empty to sync all workspaces you have access to
3. Click **Connect with Microsoft**
4. Sign in with your Microsoft account and grant access
5. Dot will start syncing your semantic models

{% hint style="success" %}
**Finding your Workspace ID**

Open your workspace in Power BI. The Workspace ID is in the URL: `https://app.powerbi.com/groups/{workspace-id}/...`
{% endhint %}

## What Gets Synced

When you connect Power BI, Dot imports:

* **Workspaces** — appear as schemas in Dot
* **Datasets / Semantic models** — appear as tables
* **Measures and columns** — with their descriptions and data types
* **Relationships** — connections between tables

Dot periodically re-syncs to pick up new or changed models. You can also trigger a manual sync from the connection settings.

## Limitations

{% hint style="warning" %}
**Known limitations**

* **Premium or PPU capacity required** — datasets on shared capacity cannot be queried via the API
* **Row-Level Security (RLS) not supported** — datasets with RLS enabled cannot be queried
* **Personal workspaces not supported** — "My Workspace" cannot be accessed via OAuth
* **Maximum 100,000 rows per query** — larger result sets are truncated
  {% endhint %}


# Azure Analysis Services

Connect Dot to Azure Analysis Services semantic models

Connect Dot to your Azure Analysis Services (AAS) server to query your semantic models using natural language. Dot syncs your tables, columns, measures, and relationships so you can ask questions about your data without writing DAX.

{% hint style="info" %}
**Requirements**

* An Azure Analysis Services server with at least one deployed model
* Read access to the models you want to query
* One of the two authentication methods described below
  {% endhint %}

## Authentication Methods

### Option 1: Interactive SSO (OAuth)

Sign in with your Microsoft account. Dot uses delegated access on your behalf — no service principal setup required.

1. Go to **Settings → Semantic Layers → Azure Analysis Services**
2. Enter your **Server URL** (e.g., `asazure://westeurope.asazure.windows.net/myserver`)
3. Optionally enter a **Database Name** to sync only that model, or leave empty to sync all models on the server
4. Click **Connect with Microsoft**
5. Sign in with your Microsoft account and grant access
6. Dot will start syncing your semantic models

{% hint style="success" %}
**Finding your Server URL**

Open your AAS resource in the Azure Portal. The server name is shown on the Overview page, formatted as: `asazure://{region}.asazure.windows.net/{servername}`
{% endhint %}

### Option 2: Service Principal

Use an Azure AD app registration for automated, unattended access — ideal for production deployments where you don't want auth tied to a user account.

**1. Create an Entra ID Application**

1. Go to **Azure Portal → Azure Active Directory → App registrations → New registration**
2. Name: `Dot-AAS-Access` (or similar)
3. Account types: Accounts in this organizational directory only
4. Click **Register**
5. Go to **Certificates & secrets → New client secret** and copy the value
6. Note the **Application (client) ID** and **Directory (tenant) ID** from the Overview page

**2. Grant AAS Access**

Add the service principal as a member of a server role:

1. Open SQL Server Management Studio (SSMS) and connect to your AAS server
2. Right-click the server → **Properties → Security**
3. Click **Add** and enter: `app:{client_id}@{tenant_id}`
4. Assign the appropriate role (Reader is sufficient for Dot)

**3. Connect in Dot**

1. Go to **Settings → Semantic Layers → Azure Analysis Services**
2. Enter your **Server URL**, **Database Name**, **Client ID**, **Client Secret**, and **Tenant ID**
3. Click **Connect**

## What Gets Synced

When you connect Azure Analysis Services, Dot imports:

* **Databases / Models** — appear as schemas in Dot
* **Tables and columns** — with their descriptions and data types
* **Measures** — including DAX expressions
* **Relationships** — connections between tables

Dot periodically re-syncs to pick up changes. You can also trigger a manual sync from the connection settings.

## Troubleshooting

### Token Expired / Please Reconnect

If you see "No access token available" or similar errors, your OAuth refresh token may have expired (tokens last 90 days). Go to **Settings → Semantic Layers → Azure Analysis Services** and click **Connect with Microsoft** again to re-authenticate.

## Limitations

{% hint style="warning" %}
**Known limitations**

* **Cross-tenant access** — querying AAS servers in a different Azure AD tenant requires B2B guest access
* **Row-Level Security (RLS)** — supported for Service Principal auth, but the SP must be added to an RLS role in SSMS
* **Maximum 100,000 rows per query** — larger result sets are truncated
  {% endhint %}


# Looker

## Create API keys

1. Go to your Looker dashboard: <https://company.cloud.looker.com>.
2. Click **Admin > Users** in the menu bar.

<figure><img src="https://files.readme.io/bddc402-1.png" alt=""><figcaption></figcaption></figure>

3. Click **Edit** next to a user.

<figure><img src="/files/iyCgvoHE2K4EUopR5YoT" alt=""><figcaption></figcaption></figure>

4. Click the **Edit Keys** button next to API Keys.

<figure><img src="/files/0U1MQb38OPTFwaE2QOO5" alt=""><figcaption></figcaption></figure>

4. Click **New API3 Key**.

<figure><img src="/files/qhFBQZVkMZrraqPPQHGi" alt=""><figcaption></figcaption></figure>

5. Copy the **Client ID** and **Client Secret**.

<figure><img src="/files/9N2daFQoSDcL4TmctAHg" alt=""><figcaption></figcaption></figure>

## Allow Dot IPs

If your organization uses a firewall to manage Looker access, Dot will only access your Looker through the following IPs:

* `5.78.211.110`
* `178.105.217.177`


# Malloy

Connect a Malloy model so Dot answers with your team's own metric definitions

Malloy is a language for modeling your data. If your team already defines dimensions and measures in Malloy, you probably don't want to write them again inside Dot. So connect your Malloy model instead. Dot reads the definitions straight from your repository, and when someone asks a question, Dot answers using your metrics rather than its own guess at what a metric means.

The point is consistency. The numbers Dot gives back match the ones your Malloy model produces, because they come from the same place. You keep one source of truth for your metrics, and you edit it where you already work, in git.

## What you need

Your Malloy model lives in a git repository. That repo holds your model files, an `index.malloy` and a `malloy-config.json`, and the config is where the warehouse connection is defined. Dot connects to the repo, reads the model, and uses that same warehouse connection, so you don't set the database up twice.

Have this ready:

* The repository, on GitHub, GitLab, or Bitbucket
* A way for Dot to read it: connect the GitHub App for one-click access, or point Dot at the repo manually
* The branch, if it isn't the default one

## Connect it

1. Sign in to Dot as an admin.
2. Open Settings, then Connections, and find the Semantic Layers section.
3. Click Malloy.
4. Connect the repository. Use Connect GitHub for one-click access, or choose Enter manually for a GitHub, GitLab, or Bitbucket repo.
5. Leave the branch blank to use the default branch, or name a specific one.
6. If your model isn't at the root of the repo, put the folder that holds `index.malloy` in Model path. Leave it blank and Dot finds it for you.
7. Click Connect.

<figure><img src="/files/3F8IRm586mpPfS3jfzbf" alt="The Malloy connection form in Dot, connecting a git repository that holds a Malloy model"><figcaption><p>Connect the repo that holds your Malloy model. Dot reads the model and the warehouse connection from it.</p></figcaption></figure>

Once the repo is connected, Dot reads the config. If the model references any secrets, like a warehouse password, Dot asks you for just those and keeps them out of your repo, filling them in when it runs a query. Then it turns your Malloy sources into tables it can query, with your dimensions and measures attached, and answers start using those definitions.

## Keeping it in sync

When you change the model in git, Dot picks up the change. You can also start a sync yourself from the connection card after you push an update.

## Row-level security

A Malloy source can carry its own access rules. If a source is filtered by group, Dot applies that filter for each person, so people only see the rows they are allowed to see. It works the same way as row-level permissions on your other tables. See [Permissions](/train-dot/permissions) for how groups and row-level filtering work.


# Steep

Connect Dot to Steep to query your metrics using natural language

Connect Dot to your [Steep](https://steep.app) semantic layer to query metrics, dimensions, and slices using natural language. Dot syncs your metric definitions, dimensions, and their descriptions so you can ask questions about your data without writing code.

{% hint style="info" %}
**Requirements**

* A **Steep account** with at least one metric defined
* A **Steep API key** — generate one in Steep under **Settings → API**
  {% endhint %}

## Connect to Steep

<figure><img src="/files/fPhp7oDIshCw4rCZIWoa" alt="Steep connection settings"><figcaption></figcaption></figure>

1. Go to **Settings → Semantic Layers → Steep**
2. Enter your **API Key**
3. Click **Connect**
4. Dot will sync your metrics and dimensions

{% hint style="success" %}
**Generating an API key**

In Steep, go to **Settings → API** and create a new API key. Copy the key and paste it into Dot.
{% endhint %}

## What Gets Synced

When you connect Steep, Dot imports:

* **Metrics** — each metric appears as a table (e.g., revenue, order\_volume)
* **Dimensions** — appear as columns on each metric, with their data types
* **Descriptions** — metric and dimension descriptions are synced for context
* **Slices** — predefined filters listed in metric descriptions (e.g., "UK", "Enterprise")
* **Related metrics** — noted in descriptions so Dot can suggest complementary data

Dot periodically re-syncs to pick up new or changed metrics. You can also trigger a manual sync from the connection settings.

## How Queries Work

When you ask a question, Dot generates a `steep_query()` call that fetches data for one metric at a time. Queries support:

* **Breakdowns** — group results by up to 2 dimensions (e.g., by Country, by Product Category)
* **Filters** — narrow results to specific dimension values
* **Slices** — use predefined filters for common segments
* **Time ranges** — specify date ranges and time granularity (daily, weekly, monthly, quarterly, yearly)


# Cube

Dot answers through your Cube semantic layer, so every number matches your governed definitions.

If your team has invested in a [Cube](https://cube.dev) semantic layer, your metric logic already lives in one governed place. Connect it to Dot and Dot won't re-derive revenue from raw tables — it asks Cube, so its answers match every other tool built on the same definitions.

## What you need

* Your Cube **REST API endpoint**, for example `https://<deployment>.cubecloud.dev/cubejs-api/v1` (or your self-hosted equivalent)
* An **API token** your Cube deployment accepts

## Connect in Dot

1. Go to **Settings → Connections** and choose **Cube**.
2. Enter the API URL and the token, then connect.

Dot syncs your cubes — measures, dimensions, and joins — into its Model, where you activate the ones Dot should use and enrich them with descriptions.

## How Dot queries Cube

Questions are answered through Cube's REST API, never with raw SQL against the underlying warehouse. That means:

* **Pre-aggregations** accelerate Dot like any other Cube client.
* The **token's security context** applies, so Cube-side access rules keep working.
* Metric math has one home: change a definition in Cube and Dot follows automatically.


# Code Repositories

Connect Dot to your code repositories


# dbt Core

Enrich your tables with dbt descriptions, SQL, lineage, and data sources

Connect a dbt project repository so Dot understands your transformation logic, lineage, and data sources. This leads to better answers -- Dot can write more accurate SQL, explain where data comes from, and answer broader questions across your data stack.

{% hint style="info" %}
**Requirements**

* A **dbt project** hosted in a Git repository (GitHub, GitLab, Bitbucket, or any HTTPS-accessible repo)
* A **database connection** already configured in Dot that matches the dbt project's target database
  {% endhint %}

## Connect Your dbt Repository

Go to **Settings** > **Connections** and scroll to find **dbt Repository**.

<figure><img src="/files/mcmFc1xeMjpLrqO9p5HG" alt="dbt Repository connection form showing Connect GitHub button, branch field, and linked database connection dropdown"><figcaption><p>The dbt Repository connection form</p></figcaption></figure>

### Option A: Via GitHub App (recommended)

1. Click **Connect GitHub** and install the Dot GitHub App on your organization
2. Select your **repository** from the dropdown
3. Set the **branch** (leave blank for the repository's default branch)
4. Select the **Linked Database Connection** -- this tells Dot which database the dbt models target
5. Click **Connect**

### Option B: Manual

1. Click **Enter manually**
2. Enter the **Repository URL** (e.g., `https://github.com/your-org/dbt-analytics`)
3. Enter an **Access Token** (optional for public repos; for private repos, use a GitHub PAT with repo read access)
4. Set the **branch** (leave blank for the repository's default branch)
5. Select the **Linked Database Connection**
6. Click **Connect**

Dot clones the repository, parses the dbt project, and matches models to your existing tables.

## What Gets Synced

Dot extracts the following from your dbt project:

* Model and column descriptions
* Model SQL
* Upstream/downstream lineage
* Root data sources
* Tags and governance metadata

Dot re-syncs periodically, or you can trigger a manual sync from the connection settings.

## What It Looks Like

Each matched table shows enriched descriptions, fields, and a compact dbt lineage indicator:

<figure><img src="/files/r1Qddfwb3UdbrCupFZs2" alt="Table drawer showing dbt-enriched descriptions, fields, and lineage indicator"><figcaption><p>A table enriched with dbt metadata -- descriptions, fields, and upstream/downstream lineage</p></figcaption></figure>

## Linked Database Connection

The **Linked Database Connection** dropdown tells Dot which database to match models against. For example, if your dbt project targets a Snowflake warehouse, select your Snowflake connection. Dot matches models by comparing the model name to table names in that connection.

## Allow Dot IPs

If your organization uses IP allowlisting to manage Git access, Dot will only access your repository through the following IPs:

* `5.78.211.110`
* `178.105.217.177`


# dbt Exposures

Publish Dot's dashboards and models back into your dbt project as exposures for lineage and impact analysis.

The [dbt Core integration](/integrations/code-repositories/dbt-core) teaches Dot about your models. This does the reverse: it exports Dot back into your **dbt project** as [exposures](https://docs.getdbt.com/docs/build/exposures) — the dbt resource that represents a downstream consumer of your models.

Commit the generated file and Dot becomes a first-class citizen of your dbt DAG. Every shared Dot dashboard — and Dot itself — shows up in `dbt docs`, `dbt ls`, and your lineage graphs, sitting downstream of the exact models it reads from.

The payoff is **impact analysis**: before you change, rename, or deprecate a model, you can see which Dot dashboards depend on it — and catch breakage in your dbt CI instead of from a confused stakeholder.

{% hint style="info" %}
**Requirements**

* A **dbt repository connected** to Dot (see [dbt Core](/integrations/code-repositories/dbt-core)) — this is what lets Dot validate every `ref()` against your project's manifest.
* **Dashboards enabled** for your workspace (exposures are derived from your shared Dot dashboards).
* An **admin or modeler** [**API token**](/developers/api). Viewer tokens are rejected with `403`.
  {% endhint %}

## What Dot exports

Dot returns a dbt v2 `exposures.yml` containing two kinds of exposures.

**One exposure per shared dashboard** (`type: dashboard`). Its `depends_on` lists the dbt models behind the dashboard's queries. Lineage is derived from each dashboard's **compiled SQL**, so the model list is exact — not guessed from names.

**One aggregate "Dot Model" exposure** (`name: dot_model`, `type: application`). Its `depends_on` is *every* active dbt model Dot has a table doc for. This represents Dot-as-a-whole as a consumer, so Dot appears in your DAG even for models that no dashboard references yet.

<figure><img src="/files/VKP18VylViQYVppaaUrY" alt="A Dot dashboard rendered as a dbt exposure in dbt docs, showing its owner, the dot_app_id and dot_tables metadata, and the orders model under Depends On"><figcaption><p>A shared Dot dashboard, rendered as an exposure in <code>dbt docs</code></p></figcaption></figure>

## Export the exposures file

Authenticate with the `X-API-KEY` header and write the response straight into your dbt project's model paths:

```bash
curl -sf -H "X-API-KEY: $DOT_API_KEY" \
  "https://app.getdot.ai/api/dbt/exposures" \
  -o models/dot_exposures.yml
```

{% hint style="info" %}
Use the host for your region: `app.getdot.ai` (US) or `eu.getdot.ai` (EU).
{% endhint %}

### Query parameters

| Parameter       | Values                   | Description                                                                                                                                      |
| --------------- | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| `format`        | `yaml` (default), `json` | `yaml` returns a ready-to-commit `exposures.yml`. `json` returns the same payload plus an `apps_without_dbt_models` list for coverage debugging. |
| `connection_id` | a dbt repo connection id | Required **only** when your org has more than one dbt repository connected. With a single repo it is selected automatically.                     |

The generated file looks like this:

```yaml
# Generated by Dot — dbt exposures declaring Dot and its dashboards as
# downstream consumers of your dbt models. Fetch from GET /api/dbt/exposures
# (X-API-KEY auth) and place under your dbt project's model-paths, e.g.
# models/dot_exposures.yml.
version: 2
exposures:
- name: dot_orders_pulse
  label: Orders Pulse
  type: dashboard
  url: https://app.getdot.ai/apps/_shared__orders-pulse
  description: 'Dot dashboard: Orders Pulse'
  owner:
    name: data-team@yourcompany.com
    email: data-team@yourcompany.com
  depends_on:
  - ref('orders')
  meta:
    dot_app_id: _shared__orders-pulse
    dot_tables:
    - analytics.public.orders
- name: dot_model
  label: Dot Model
  type: application
  url: https://app.getdot.ai/apps
  description: Dot AI data analyst — downstream consumer of every dbt model Dot has table documentation for. Changes to these models can affect Dot's answers, dashboards, and scheduled reports.
  owner:
    name: Dot
  depends_on:
  - ref('customers')
  - ref('orders')
  meta:
    dot_aggregate: true
    dot_model_count: 2
```

## Add it to your dbt project

Drop the file anywhere under your configured `model-paths` (e.g. `models/dot_exposures.yml`), validate, and commit:

```bash
dbt parse   # validates every ref() resolves
git add models/dot_exposures.yml
git commit -m "Add Dot exposures"
```

Run `dbt docs generate` and Dot's dashboards appear alongside your models.

## Impact analysis

This is where exposures earn their keep. Once the file is committed, dbt's own selectors reveal Dot's dependence on any model:

```bash
# Which exposures — Dot dashboards and Dot itself — depend on fct_orders?
dbt ls --select fct_orders+ --resource-type exposure
```

The same relationship shows up visually in the lineage graph, where each Dot dashboard sits at the downstream end of the DAG:

<figure><img src="/files/pto2V0YMeEVbfcwijQZ9" alt="A dbt lineage graph: raw_orders and raw_payments flow through staging models into the orders model, which feeds the Orders Pulse Dot exposure at the end of the DAG"><figcaption><p>A Dot dashboard as the downstream endpoint of a dbt lineage graph (<code>+exposure:dot_orders_pulse</code>)</p></figcaption></figure>

Wire this into your dbt CI and a pull request that touches an upstream model will surface exactly which Dot dashboards it puts at risk — before it merges.

## Keep it in sync

Dashboards and models change, so re-export on a schedule to keep the file current. A minimal GitHub Actions job:

```yaml
# .github/workflows/dot-exposures.yml
name: Sync Dot exposures
on:
  schedule:
    - cron: "0 6 * * 1-5"   # weekday mornings
  workflow_dispatch:
jobs:
  sync:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Fetch exposures from Dot
        run: |
          curl -sf -H "X-API-KEY: ${{ secrets.DOT_API_KEY }}" \
            "https://app.getdot.ai/api/dbt/exposures" \
            -o models/dot_exposures.yml
      - name: Validate
        run: dbt parse
      - name: Commit if changed
        run: |
          git config user.name  "dot-bot"
          git config user.email "dot-bot@users.noreply.github.com"
          git add models/dot_exposures.yml
          git diff --cached --quiet || git commit -m "chore: sync Dot exposures"
          git push
```

## Safe by construction

The export is designed to never break your dbt project:

* **Every `ref()` is validated** against the dbt manifest captured at sync time. A model that isn't in the manifest — renamed, disabled, or versioned — is **withheld** into `meta.dot_unresolved_models` on the affected exposure rather than emitted, so `dbt parse` always succeeds. (Tables that match more than one Dot table doc are likewise reported under `meta.dot_ambiguous_tables` instead of being guessed.)
* **Deterministic output** — no timestamps or run-specific ordering, so re-exports diff clean in your repo and only change when your dashboards or models actually change.
* **Nothing is silently dropped.** Dashboards built only on non-dbt or data-only sources can't produce `ref()`s; request `?format=json` to see them listed under `apps_without_dbt_models`.

## Surface data-quality incidents on dashboards

Exposures tell dbt what Dot depends on. The reverse channel lets dbt (and other tools) tell **Dot** when that upstream data is currently failing — so viewers see it in context.

{% hint style="info" %}
Incident ingest writes org-wide dashboard indicators, so `POST /api/dbt/run_results` and `POST /api/quality/incidents` require an **admin API token**. Modeler and viewer tokens are rejected.
{% endhint %}

After a `dbt build`, send Dot your run results together with the manifest. Dot opens incidents for failing models and tests and stale sources, and clears them automatically when they pass again or leave the manifest:

```bash
dbt build   # writes target/run_results.json and target/manifest.json

jq -n \
  --slurpfile run target/run_results.json \
  --slurpfile manifest target/manifest.json \
  '{run_results: $run[0], manifest: $manifest[0]}' > dot_run.json

curl -sf -X POST "https://app.getdot.ai/api/dbt/run_results" \
  -H "X-API-KEY: $DOT_API_KEY" \
  -H "Content-Type: application/json" \
  --data @dot_run.json
```

Not on dbt? `POST /api/quality/incidents` accepts incidents from any source (Airflow, Monte Carlo, your own checks):

```bash
curl -sf -X POST "https://app.getdot.ai/api/quality/incidents" \
  -H "X-API-KEY: $DOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "source": "airflow",
        "items": [
          {
            "table": "analytics.public.orders",
            "check": "freshness",
            "status": "open",
            "severity": "warning",
            "message": "orders load is 6h late",
            "url": "https://airflow.yourcompany.com/dags/orders"
          }
        ]
      }'
```

Any Dot dashboard built on an affected table then shows a subtle **"N data issues"** indicator in its top bar, with the details on hover — so viewers know upstream data is currently failing before they trust the numbers.


# BI Tools

Connect your BI tools so Dot learns your business logic, and rebuild dashboards as Dot apps

Connect the BI tools your team already uses, like Tableau, Metabase, Sigma, and Qlik. Dot reads your dashboards to learn how your business defines its metrics. That way its answers match the numbers people already trust.

Once a BI tool is connected, there's a second thing you can do: rebuild one of its dashboards as a Dot app.

## Migrate a dashboard to Dot

You can turn a Tableau, Metabase, or Sigma dashboard into a Dot app. Ask Dot to migrate a dashboard by name and it recreates it piece by piece, the same charts and tables and layout, but now backed by live SQL you can question and change in plain language.

Why do this? Some teams want to move off their BI tool and keep the dashboards they rely on. Others just want a live, conversational version sitting next to the original, so anyone can drill in without opening the BI tool.

Dot checks its own work as it goes. For each tile it compares its result against the original and marks whether the two match. When every tile matches, the migration is a true copy of the dashboard. If a tile can't be matched exactly, Dot tells you which one, so you can decide whether the difference matters.

To try it, connect the BI tool first, then ask Dot something like "migrate our Weekly Revenue dashboard from Tableau."

## Connect a BI tool

* [Tableau](/integrations/bi-tools/tableau)
* [Metabase](/integrations/bi-tools/metabase)
* [Sigma](/integrations/bi-tools/sigma)
* [Qlik](/integrations/bi-tools/qlik)


# Tableau

Dot supports Tableau Cloud and Tableau Server 2019.3 or later. Connect Tableau in two stages:

| Connection        | Required | Used for                                                                              |
| ----------------- | -------- | ------------------------------------------------------------------------------------- |
| **Standard**      | Yes      | Syncing workbooks, views and data sources; adding metadata and lineage when available |
| **Connected App** | Optional | Reading the exact values shown in dashboard tiles                                     |

## Standard connection

The Standard connection uses a Personal Access Token (PAT). Set this up first for both Tableau Cloud and Tableau Server.

### 1. Create a Personal Access Token in Tableau

Use a Tableau user with the **Site Admin Explorer** or **Site Admin Creator** role.

1. Open **My Account Settings**.

![Tableau account menu with My Account Settings](https://docs.sled.so/~gitbook/image?url=https%3A%2F%2F2457798860-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FEdkjXblQVFxdE5uvkTZt%252Fuploads%252Fbqm2rhWu90uUXeiDhJAW%252Fgrafik.png%3Falt%3Dmedia%26token%3D42f9c7de-8267-43c5-af2c-ec953084ad46\&width=300\&dpr=4\&quality=100\&sign=606aa2ae\&sv=2)

2. Find **Personal Access Tokens**.

![Personal Access Tokens section in Tableau](https://docs.sled.so/~gitbook/image?url=https%3A%2F%2F2457798860-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FEdkjXblQVFxdE5uvkTZt%252Fuploads%252FLlV9TCmQmJBK3a5OlfVM%252Fgrafik.png%3Falt%3Dmedia%26token%3Dca573f80-d452-418d-9ebb-bbf5bb8e4427\&width=768\&dpr=4\&quality=100\&sign=bd2ab74e\&sv=2)

3. Enter a name and select **Create new token**.

![Create a new Personal Access Token in Tableau](https://docs.sled.so/~gitbook/image?url=https%3A%2F%2F2457798860-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FEdkjXblQVFxdE5uvkTZt%252Fuploads%252FeZwVf7mYcTPLPKUySC3S%252Fgrafik.png%3Falt%3Dmedia%26token%3Dab3c7dd0-e124-4aee-aa4d-936268c67121\&width=300\&dpr=4\&quality=100\&sign=3beb599f\&sv=2)

4. Copy the token name and secret and store the secret securely.

### 2. Connect in Dot

Open the Tableau connection in Dot and stay on the **Standard** tab. Enter:

* **Server URL**
* **Site ID**, if you do not use the default site
* **Token Name**
* **Token Value**

Select **Connect** and wait for the first sync to finish. You can then curate the synced Tableau content under **Model → External assets**.

<figure><img src="/files/SuFx5OGYMiKPEaMwa4FU" alt="Standard Tableau connection fields in Dot"><figcaption></figcaption></figure>

## Connected App

Add a Connected App after the Standard connection if you want Dot to read the exact values shown in Tableau dashboard tiles, including the active filters and period.

### 1. Create a Direct Trust Connected App in Tableau

1. Go to **Settings → Connected Apps** and create an app of type **Direct Trust**, not OAuth 2.0. On Tableau Server, the UI is available in version 2022.1 or later; version 2021.4 supports connected apps through the REST API only.
2. Add the Dot host where your workspace runs to the app's domain allowlist.
3. Enable the Connected App. Tableau creates new apps in a disabled state.
4. Generate a secret and copy the **Client ID**, **Secret ID** and **Secret Value**.

See Tableau's guide to [configuring a Direct Trust Connected App](https://help.tableau.com/current/server/en-gb/connected_apps_direct.htm).

### 2. Connect it in Dot

Open the existing Tableau connection and select the **Connected App** tab. Enter:

* **Client ID**
* **Secret ID**
* **Secret Value**
* **Tableau admin user**

<figure><img src="/files/vWcySQWhmB52DuWwTYFY" alt="Connected App fields in the Tableau connection in Dot"><figcaption></figcaption></figure>

Use an existing licensed Tableau administrator who can access the workbooks Dot should check. Dot uses this account for automated checks. Other users keep their own Tableau permissions.

Select **Save Connected App**. Dot checks the configuration and shows when it was verified or why it failed. Use **Run check again** to retest it later without re-entering the credentials.

## Tableau Server setup

Tableau Cloud manages the services and network access described below. If you use Tableau Server, review the requirements for each connection.

### Standard connection

#### Network access

Dot must be able to reach the Tableau Server URL. Allow Dot's service IPs, `5.78.211.110` and `178.105.217.177`, to access `/api/*`.

If the server is not reachable from the internet, coordinate with our customer success team at <hi@getdot.ai>. A typical private-network setup uses OpenVPN.

#### Metadata API for full lineage

The Standard connection works without the Metadata API: Dot can still list workbooks and views. Enable the Metadata API if you want Dot to also sync lineage to warehouse tables and calculated-field definitions.

The Metadata API is installed and disabled by default on Tableau Server. The **Tableau Catalog** checkbox in the site UI is separate and does not confirm that the server-level Metadata API is running.

Ask a Tableau Server administrator to verify the service from the initial server node:

```bash
tsm maintenance metadata-services get-status
```

If Dot receives `403 Forbidden` from `/relationship-service-war/graphql`, the server is reachable but the Metadata API is not enabled.

If the Metadata API is not running or its store is not initialized, enable it with:

```bash
tsm maintenance metadata-services enable
```

Enabling it starts metadata indexing and temporarily restarts some Tableau services. See [Tableau's Metadata API guide](https://www.tableau.com/developer/learning/metadata-api#tab-325797-0).

### Connected App

Dot reads exact values through Tableau's Embedding API in a server-side browser. Allow the same Dot service IPs to reach these additional paths:

* `/auth/*`
* `/javascripts/*`
* `/views/*`
* `/vizql/*`
* `/vizportal/*`

If an identity-aware proxy or corporate SSO gateway sits in front of Tableau, it may redirect those browser requests to an interactive login page before they reach Tableau. In that case, the Standard connection can sync successfully while the Connected App check fails.

Ask whoever manages the proxy to exempt Dot's service IPs for the paths above. When Dot detects a redirect, the connection check names the identity provider that intercepted it.


# Metabase

Both Metabase Cloud or self-hosted Metabase are supported.

To connect Dot to Metabase, follow these steps:

## **1. Generate an API Key in Metabase**

Metabase allows the creation of API keys to authenticate programmatic requests. To generate an API key:

1. Click on the gear icon in the upper right corner of Metabase.
2. Select **Admin settings**.
   1. ![](/files/JaMePreYC2aLkLQEhWgT)
3. Navigate to the **Settings** tab.
4. Click on the **Authentication** tab in the left-hand menu.
5. Scroll to the **API Keys** section and click **Manage**.
6. Click the **Create API Key** button. 1.

   ```
   <figure><img src="../../../.gitbook/assets/image (12).png" alt=""><figcaption></figcaption></figure>
   ```
7. Enter a descriptive **Key name** to identify its purpose.
8. Select a **Group** to assign the key, determining its permissions.
9. Click **Create**.
10. Copy the generated API key and store it securely, as Metabase will not display it again.

## **2. Obtain the Metabase Server URL**

Identify the base URL of your Metabase instance. This is typically in the format `http://your-metabase-domain.com` or `https://your-metabase-domain.com`.

## **3. Connect Metabase to Dot**

In Dot, provide the Metabase Server URL and the API key you generated:

1. Open Dot and navigate to **Settings** / **Connections**.
2. Select the option to connect to Metabase.
3. Enter the **Server URL** (your Metabase base URL).
4. Input the **API Key** obtained from Metabase.
5. Click **Connect** to establish the integration.

<figure><img src="/files/UeCro0NzvmToybZll6nk" alt=""><figcaption></figcaption></figure>

Once connected, Dot will synchronize with Metabase. As soon as it's done, you can head over to **Model** / **External assets** to further curate what Dot should know about.

**Note:** Ensure that the API key has appropriate permissions assigned through its associated group in Metabase to allow Dot to access the necessary data.


# Sigma

**Create an Client ID and Client Secret**

To create a Sigma access token for Dot visit Sigma and log into your account. Use the following steps to generate an access token:

1. Click on your avatar in the top right and select 'Administration' from the dropdown menu.
2. Click the option **APIs and Embed Secrets**
3. Click **Create New** in the top right

<figure><img src="/files/cGPskyBwOIGlVeyn7vn1" alt=""><figcaption></figcaption></figure>

1. Fill out the form with the Name and Owner. Make sure that the owner is an Administrator.
2. Copy the Client ID and Secret and save it. This will be used to connect Sigma to Dot.

**Retrieve your Sigma host**

In the **Administrator** settings go to **Account.** In the top under the **Site** you can see which cloud you are hosted on.

<figure><img src="/files/4lSdKo6S2ewa3WvIEhqu" alt=""><figcaption></figcaption></figure>

You can find your respective host parameter in [this table](https://help.sigmacomputing.com/reference/get-started-sigma-api#identify-your-api-request-url). For example, if you are hosted on AWS US (West), your host parameter is <https://aws-api.sigmacomputing.com>.


# Qlik

Connect Qlik so Dot can learn from your Qlik apps

Connect your Qlik Cloud tenant and Dot can read your Qlik apps. It uses them the same way it uses your other BI tools: to learn which metrics your team trusts and how your dashboards define them, so its own answers line up with what people already look at.

## What you need

In your Qlik Cloud tenant, create an OAuth client (Qlik calls this a machine-to-machine OAuth client). From it, get these three values:

* Your Qlik server URL, for example `https://your-tenant.qlikcloud.com`
* The Client ID
* The Client Secret

## Connect it

1. Sign in to Dot as an admin.
2. Open Settings, then Connections, and find the BI Tools section.
3. Click Qlik.
4. Enter your server URL, Client ID, and Client Secret.
5. Click Connect.

<figure><img src="/files/0SOY8ZM1c90Jb4smYnNy" alt="The Qlik connection form in Dot, with Server URL, Client ID, and Client Secret fields"><figcaption><p>Connect Qlik with your server URL and OAuth client credentials.</p></figcaption></figure>

Once connected, Dot syncs your Qlik apps and can reference them when it answers questions and when it builds your data model. To see what it picked up, head over to the Model page.


# Knowledge Bases

Bring your team's written knowledge — docs, wikis, and tickets — into Dot so its agents can search, read, ask, and (where enabled) write.

Most of what a data team needs to answer a question well isn't in the warehouse — it's the tribal knowledge written down in a wiki, a notes app, or a ticket. Knowledge-base connectors let Dot reach that content directly: its agents can **search**, **read**, and **ask questions** across your documentation, and for some tools **create or update** pages and issues on your behalf.

These are the connectors that answer the original ask behind this whole category — *"all our company documentation lives in Slite / Confluence / Notion, how do I feed that to Dot?"*

## Available connectors

| Connector                                                  | Good for                | Dot can…                                                                                          |
| ---------------------------------------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------- |
| [Slite](/integrations/knowledge-bases/slite)               | Slite workspaces        | Ask (hosted Q\&A), search, read notes — **read-only**                                             |
| [Notion](/integrations/knowledge-bases/notion)             | Notion workspaces       | Search, read pages & databases, create/update pages, comments                                     |
| [Confluence](/integrations/knowledge-bases/confluence)     | Atlassian Cloud wikis   | Search, read, create/update pages, attachments, comments                                          |
| [Jira](/integrations/knowledge-bases/jira)                 | Atlassian Cloud issues  | Search, read, create/update/transition issues, comments                                           |
| [Google Drive](/integrations/knowledge-bases/google-drive) | Google Workspace Drives | Search and read Docs, Sheets, Slides, PDFs — **read-only**, scoped per shared drive (or per user) |

## How they work

**1. Connect once, in Settings.** Each connector is a card under **Settings → Connections**, in the group the app labels **Context Connectors**. You supply an API token — or authorize via OAuth for Notion, or a service account for Google Drive — and Dot validates it.

**2. Govern each action, in Model → Skills.** Once connected, the connector appears as a **skill** under **Model → Skills**. Expand **Configure permissions** to see each action as its own toggle — an admin turns individual actions on or off. A disabled action is refused with a message naming the exact permission to flip, so nothing runs that you haven't allowed.

**3. Dot uses it at query time.** Unlike a database sync, these connectors don't copy your content into Dot. Dot's agents reach into the source *when a question needs it* — searching, reading, or asking against the live workspace. This keeps answers current and means access always reflects live permissions in the source tool — those of the connected account, or for [Google Drive](/integrations/knowledge-bases/google-drive), of whichever drive that connection was scoped to.

{% hint style="info" %}
**Permission defaults follow a simple rule.** Non-destructive reads (and, for Notion/Confluence/Jira, ordinary writes) default **on** — connecting the tool implies you want Dot to use it. Destructive or privacy-sensitive actions (delete, archive, listing workspace members) default **off** and must be turned on explicitly.
{% endhint %}

## Related

Knowledge-base connectors are one of the tools available to [Root, Dot's Context Agent](/train-dot/context-agent) — the agent that builds and maintains your organization's knowledge base. Pointing Root at your Confluence space or Slite workspace lets it extract business logic straight from your existing documentation.


# Slite

Ask questions over your Slite knowledge base and let Dot search and read notes.

Connect your [Slite](https://slite.com) workspace so Dot can answer questions from your company documentation. Slite exposes its own hosted question-answering endpoint, so the fastest path — *"what does our Slite say about X?"* — is a single **ask** call that returns a synthesized answer with cited sources.

{% hint style="info" %}
**Slite is read-only today.** Dot can ask, search, and read notes. It cannot yet create, edit, archive, or delete Slite content — if you ask it to, it will tell you writing isn't supported rather than attempt a workaround.
{% endhint %}

## Prerequisites

* A Slite workspace and a Dot admin account.
* A Slite **API key**. Any workspace member can generate a personal key; it inherits that member's access, so use an account that can see the docs you want Dot to reach.

## Get a Slite API key

1. In Slite, open **Settings → API**.
2. Click **Create a new key**.
3. Copy the key — you'll paste it into Dot next.

## Connect in Dot

1. Go to **Settings → Connections** and find **Slite** under **Context Connectors**.
2. Paste your API key into the **API key** field.
3. Click **Connect Slite**. Dot validates the key against your workspace and confirms the connection.

## What Dot can do

Each action is a separately governed permission under **Model → Skills → Slite**. All three default **on** — connecting Slite implies you want agent-driven reads.

| Permission         | What it allows                                                                                                         |
| ------------------ | ---------------------------------------------------------------------------------------------------------------------- |
| `slite.ask`        | Natural-language Q\&A over the whole workspace, with cited sources. The best answer to "what do our docs say about X". |
| `slite.search`     | Keyword search across notes.                                                                                           |
| `slite.notes.read` | Reading a note's content and navigating the note tree.                                                                 |

To turn one off, open **Model → Skills**, expand the **Slite** skill, and toggle the permission. A disabled action is refused with a message naming the permission an admin must re-enable.

## Limitations

* **Read-only.** No create / update / delete / archive path yet.
* Dot reads Slite live at query time — it does not sync or copy notes into Dot, so answers always reflect the current workspace and the connected key's access.


# Notion

Let Dot read and write your Notion pages, databases, and comments.

Connect your [Notion](https://notion.com) workspace so Dot can search, read pages and databases, and — when enabled — create or update pages and post comments. Writes are attributed to Dot: agent-created page titles are prefixed with `[Dot]`, and comments carry a `[Dot, on behalf of <user>]` byline so authorship stays traceable.

## Prerequisites

* A Notion workspace and a Dot admin account.
* Permission to authorize an integration and choose which pages it may access.

## Connect in Dot

1. Go to **Settings → Connections** and find **Notion** under **Context Connectors**.
2. Click **Connect Notion**. This opens Notion's authorize page.
3. Choose the pages and databases the integration may access, and confirm. Dot shows the connected workspace once it's done.

{% hint style="warning" %}
**Notion shares by page.** The integration only sees pages you explicitly share with it. If Dot's searches come back empty, share more pages (or their parent) with the integration in Notion.
{% endhint %}

## What Dot can do

Each action is a separately governed permission under **Model → Skills → Notion**.

| Permission              | What it allows                                      | Default |
| ----------------------- | --------------------------------------------------- | ------- |
| `notion.search`         | Title search across reachable pages and databases   | On      |
| `notion.pages.read`     | Reading pages and their block content               | On      |
| `notion.pages.write`    | Creating and updating pages, appending content      | On      |
| `notion.pages.archive`  | Archiving (soft-deleting) or restoring pages        | Off     |
| `notion.databases.read` | Reading database schemas and querying rows          | On      |
| `notion.comments.read`  | Reading comment threads                             | Off     |
| `notion.comments.write` | Posting comments                                    | Off     |
| `notion.users.read`     | Listing workspace members (surfaces names & emails) | Off     |

To change a permission, open **Model → Skills**, expand the **Notion** skill, and toggle it.

{% hint style="info" %}
**Comments and member listing need a second opt-in on Notion's side.** Those permissions also depend on capabilities configured on the Notion integration itself (e.g. "Read comments"). That's why they default off — enable the capability in Notion's integration settings *and* the matching permission in Dot. Turning on only one side still refuses the call.
{% endhint %}

## Limitations

* The integration sees only pages shared with it (see above).
* Markdown ↔ Notion conversion covers common blocks (paragraphs, headings, lists, to-dos, code, quotes, dividers, links, inline formatting). Toggle, callout, and synced blocks degrade to plain paragraphs.
* Very large or deeply nested pages are read in bounded chunks.


# Confluence

Let Dot read and write your Confluence pages, spaces, and comments.

Connect your Atlassian Cloud account so Dot can search, read, and — when enabled — create or update Confluence pages, upload attachments, and post comments. Dot writes plain Markdown and converts it to Confluence's native format, so tables, headings, lists, and embedded charts render correctly.

{% hint style="info" %}
**One connection, two skills.** Confluence and [Jira](/integrations/knowledge-bases/jira) share a single **Atlassian (Confluence + Jira)** connection and API token, but each is a separate skill you toggle independently under **Model → Skills**. Connecting once enables both.
{% endhint %}

## Prerequisites

* An **Atlassian Cloud** site (Server / Data Center is not supported).
* A Dot admin account and an Atlassian **API token**.

## Get an Atlassian API token

1. Go to [id.atlassian.com/manage-profile/security/api-tokens](https://id.atlassian.com/manage-profile/security/api-tokens).
2. Click **Create API token**, name it, and copy the value.

The token acts as the account that created it — Confluence records that account as the author of pages and comments Dot writes.

## Connect in Dot

1. Go to **Settings → Connections** and find **Atlassian (Confluence + Jira)** under **Context Connectors**.
2. Fill in:
   * **Site URL** — e.g. `https://your-org.atlassian.net`
   * **Account email** — the email of the account that owns the token
   * **API token** — the token you just created
3. Click **Connect**.

## What Dot can do

Each action is a separately governed permission under **Model → Skills → Confluence**.

| Permission                  | What it allows                                 | Default |
| --------------------------- | ---------------------------------------------- | ------- |
| `confluence.search`         | Cross-space full-text search                   | On      |
| `confluence.spaces.read`    | Listing spaces                                 | On      |
| `confluence.pages.read`     | Reading pages, ancestors, and page trees       | On      |
| `confluence.pages.write`    | Creating/updating pages, uploading attachments | On      |
| `confluence.pages.delete`   | Trashing pages (reversible for 30 days)        | Off     |
| `confluence.comments.read`  | Reading footer and inline comments             | On      |
| `confluence.comments.write` | Posting footer comments                        | On      |

To change a permission, open **Model → Skills**, expand the **Confluence** skill, and toggle it.

## Attribution

Pages and comments Dot writes are marked as its own: page titles are prefixed with `[Dot]`, page bodies end with a *"…on behalf of \<user>"* byline, and comments are prefixed with `[Dot, on behalf of <user>]`.

## Limitations

* **Atlassian Cloud only** — no Server / Data Center.
* Markdown conversion covers paragraphs, headings, lists, GFM tables, quotes, code blocks, dividers, links, and inline formatting; images embed via attachment upload. Confluence-specific macros (panel, expand, status) render as plain paragraphs.


# Jira

Let Dot read, create, and transition your Jira issues.

Connect your Atlassian Cloud account so Dot can search issues with JQL, read them, and — when enabled — create, edit, transition, or comment on issues. A common pattern is letting Dot file a ticket straight from an analysis: *"usage dropped for this customer — open a Jira issue with the findings."*

{% hint style="info" %}
**One connection, two skills.** Jira and [Confluence](/integrations/knowledge-bases/confluence) share a single **Atlassian (Confluence + Jira)** connection and API token, but each is a separate skill you toggle independently under **Model → Skills**. Connecting once enables both.
{% endhint %}

## Prerequisites

* An **Atlassian Cloud** site (Server / Data Center is not supported).
* A Dot admin account and an Atlassian **API token**.

## Get an Atlassian API token

1. Go to [id.atlassian.com/manage-profile/security/api-tokens](https://id.atlassian.com/manage-profile/security/api-tokens).
2. Click **Create API token**, name it, and copy the value.

The token acts as the account that created it — Jira records that account as the reporter of issues and author of comments Dot writes.

## Connect in Dot

If you've already connected Atlassian for Confluence, Jira is ready — skip this. Otherwise:

1. Go to **Settings → Connections** and find **Atlassian (Confluence + Jira)** under **Context Connectors**.
2. Fill in **Site URL** (e.g. `https://your-org.atlassian.net`), **Account email**, and **API token**.
3. Click **Connect**.

## What Dot can do

Each action is a separately governed permission under **Model → Skills → Jira**.

| Permission               | What it allows                                 | Default |
| ------------------------ | ---------------------------------------------- | ------- |
| `jira.search`            | JQL search across issues                       | On      |
| `jira.issues.read`       | Reading issues and their available transitions | On      |
| `jira.issues.write`      | Creating/editing issues, uploading attachments | On      |
| `jira.issues.transition` | Moving an issue to a new status                | On      |
| `jira.issues.delete`     | Deleting issues (not recoverable)              | Off     |
| `jira.comments.read`     | Reading issue comments                         | On      |
| `jira.comments.write`    | Adding issue comments                          | On      |

To change a permission, open **Model → Skills**, expand the **Jira** skill, and toggle it.

## Attribution

Issues and comments Dot writes are marked as its own: issue summaries are prefixed with `[Dot]`, descriptions end with a *"…on behalf of \<user>"* byline, and comments are prefixed with `[Dot, on behalf of <user>]`.

## Limitations

* **Atlassian Cloud only** — no Server / Data Center.
* Descriptions and comments accept Markdown (paragraphs, headings, lists, GFM tables, quotes, code blocks, links, inline formatting), converted to Jira's format on submission. Attachments upload to the issue's Attachments panel; inline-embedding an image in a description isn't supported.


# Google Drive

Let Dot search and read your Google Drive — Docs, Sheets, Slides, and PDFs — scoped to exactly the folders or drives you choose.

Connect [Google Drive](https://drive.google.com) so Dot can search across your documents and read them while answering questions — the pricing doc, the runbook, the QBR deck, the spec nobody can find.

The connector is **read-only**. Dot never creates, edits, or deletes anything in Drive.

## Two ways to scope access

You choose per connection, when you connect. Most workspaces want the first.

### A folder or shared drive (default)

You share **one folder** (or a shared drive) with a service account, and everyone you grant access to sees it. The credential can reach nothing else — not because Dot filters it, but because Google never gave it access to anything else. Sharing applies to everything inside the folder, so subfolders come along automatically.

Different teams get different content: repeat the setup with a **separate service account per scope**. Each becomes its own connection with its own Access Groups, so Finance sees the Finance folder and nobody else's.

{% hint style="info" %}
**Why this is the default.** Dot chooses which credential to use; Google decides what that credential can see. Dot keeps no copy of your Drive permissions, so nothing can drift out of sync with them. It needs no Workspace admin — sharing a folder is something anyone who owns one can do, on every Workspace edition — and it works in Slack and Teams.
{% endhint %}

### Each user's own Drive

Every request is made **as the person asking**, so each user sees exactly the files they could open in a browser. The strongest confidentiality, and the right answer for some companies.

What it costs: a Workspace **super-admin** must authorize the service account, that key can then read any file in the domain, and it can't be used from Slack or Teams — those run under a shared bot identity, and this mode needs a specific person.

## Prerequisites

* **Google Workspace** — personal Gmail accounts aren't supported.
* A **Google Cloud project** to hold a service account.
* A **Dot admin** account.
* For per-user mode only: a **Google Workspace super-admin**.

## Set up

### 1. Create a service account (both modes)

1. In [Google Cloud → Service accounts](https://console.cloud.google.com/iam-admin/serviceaccounts), create a service account and download a **JSON key**.
2. Enable the [Google Drive API](https://console.cloud.google.com/apis/library/drive.googleapis.com) on the same project.

### 2a. Folder or shared drive mode

In Google Drive, right-click the folder you want Dot to read → **Share**, and add the service account's **`client_email`** as a **Viewer**. Untick *Notify people* — a service account has no inbox. If your edition has shared drives, adding it as a member of one works the same way.

What you shared *is* the scope — there is nothing else to configure, and nothing else is reachable.

{% hint style="info" %}
**Scope stays live.** Sharing another folder with the same service account adds it to this connection straight away — no reconnect, no admin step. That is deliberate: it lets a team add material to Dot without filing a ticket. The flip side is that anyone who can share a folder can widen what the connection reaches, so treat the service account's address as something to hand out carefully. If you need a scope that cannot grow that way, use a shared drive — its membership is managed.
{% endhint %}

### 2b. Per-user mode

In the [Admin console](https://admin.google.com/ac/owl/domainwidedelegation) (**Security → Access and data control → API controls → Domain-wide delegation**), click **Add new**:

| Field        | Value                                            |
| ------------ | ------------------------------------------------ |
| Client ID    | the service account's numeric Unique ID          |
| OAuth scopes | `https://www.googleapis.com/auth/drive.readonly` |

{% hint style="warning" %}
**This entry is the boundary, not the key file.** The service account can only ever do what you authorize here. Grant the read-only Drive scope and nothing else.
{% endhint %}

### 3. Connect in Dot

Go to **Settings → Connections → Google Drive**, pick the mode, paste the key, and click **Connect**.

<figure><img src="/files/voNk69JH8iGOOkVEwhN3" alt="The Google Drive connection form in Dot, showing the choice between one shared drive and each user&#x27;s own Drive, with the service account key field and setup steps"><figcaption><p>The setup steps change with the mode you pick.</p></figcaption></figure>

Dot verifies against Google **before saving**. If the account can't see anything yet, the connect is refused and the error names the exact address to share with — so a half-finished setup can't sit there looking connected and then fail for everyone later.

<figure><img src="/files/EOdB8knlzHrUhDnPIi8a" alt="A connected Google Drive card in Dot, showing the folder it is scoped to and the access groups that may use it"><figcaption><p>The card names the scope, and Access Groups decides who can use it.</p></figcaption></figure>

### 4. Choose who can use it

Set **Access Groups** on the connection. Only users in those groups can reach that content through Dot.

## Using it

Just ask. Dot decides when Drive is relevant:

* *"Search Drive for the Q3 pricing doc and summarise it."*
* *"Which documents mention the Acme migration?"*
* *"What changed in the onboarding runbook recently?"*

Dot searches **Drive's own full-text index**, which covers the contents of Docs, Sheets, Slides and PDFs — not just file names. If you can reach several drives, one search covers all of them and each result says which drive it came from.

<figure><img src="/files/1dCxRQG8pRxnBOGQmMJN" alt="Dot answering a question from a document in Drive, with a link back to the source file"><figcaption><p>Answers cite the file they came from.</p></figcaption></figure>

## What Dot can do

Each action is a separately governed permission under **Model → Skills → Drive**.

| Permission               | What it allows                                   | Default |
| ------------------------ | ------------------------------------------------ | ------- |
| `drive.search`           | Full-text search across Drive contents           | On      |
| `drive.files.read`       | Listing folders, file metadata, reading contents | On      |
| `drive.files.download`   | Downloading a file to work with it               | On      |
| `drive.permissions.read` | Seeing who a file is shared with                 | Off     |

Each can also be scoped to specific user groups.

{% hint style="info" %}
**Why sharing visibility defaults off.** *"Who else can see this file"* is a different and more sensitive question than *"what does this document say"* — useful for governance reviews, but not something to hand out by default.
{% endhint %}

## How Dot reads your files

Google Docs, Sheets, and Slides have no downloadable text of their own, so Dot converts them on the fly:

| File type                 | Dot reads it as                           |
| ------------------------- | ----------------------------------------- |
| Google Doc                | Markdown                                  |
| Google Slides             | Plain text                                |
| Google Sheet              | CSV — **first tab only**                  |
| Text, Markdown, CSV, JSON | As-is                                     |
| PDF, images, Office files | Downloaded, then read with the right tool |

Long documents are read in pages, so Dot can work through a large file without losing the thread.

{% hint style="warning" %}
**For spreadsheet data, use the Google Sheets connection instead.** Drive reads a sheet's first tab as text, which is fine for context but not for analysis. To filter, aggregate, or join spreadsheet data, connect it as a [data source](/integrations/databases) so Dot can query it properly.
{% endhint %}

## Limitations

* **Read-only.** Dot cannot create, edit, or delete Drive files.
* **Google Workspace only** — personal Gmail accounts aren't supported.
* **A connection covers everything shared with its service account.** To scope more narrowly, share less — or use a second service account.
* **Sheets are read one tab at a time** (see above).
* **Per-user mode only:** unavailable from Slack and Teams, and a Dot user whose email isn't a Google account in your domain can't use it.
* Very large files are read up to a bounded size, and very large downloads are refused rather than truncated.

## Troubleshooting

| Message                                            | What it means                                                                                                    | What to do                                                                 |
| -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- |
| *can't see anything yet*                           | Nothing has been shared with the service account                                                                 | Share the folder with its `client_email` as a Viewer                       |
| *is a member of N shared drives*                   | One account was added to several drives                                                                          | Give each drive its own service account                                    |
| *Google would not issue Drive access for \<email>* | Per-user mode: the Admin console authorization is missing, or that address isn't a Google account in your domain | Recheck the client ID and scope                                            |
| *Action requires permission `drive.…`*             | That action is switched off for this user                                                                        | Enable it under **Model → Skills → Drive**                                 |
| *Google denied access to this file*                | The credential genuinely can't read it                                                                           | Share the file with the service account, or with the user in per-user mode |
| *Drive is unavailable in Slack and Teams*          | Per-user mode has no individual to act as                                                                        | Use a shared-drive connection, which works there                           |

## Related

* [Knowledge Bases overview](/integrations/knowledge-bases) — how connectors like this one fit together
* [Root, Dot's Context Agent](/train-dot/context-agent) — can use Drive to extract business logic from your existing documents


# Slack & Teams

Use Dot where your team works


# Slack

Use Dot where you get work done

<figure><picture><source srcset="/files/GKXZckYKXCNZ1a6kmtyC" media="(prefers-color-scheme: dark)"><img src="/files/7fpXlQ5UOZdlNmcIcgo6" alt=""></picture><figcaption></figcaption></figure>

## Add @Dot to Slack

1. Click on Slack in [Settings](https://app.getdot.ai/settings)

![](/files/ooncblaXK2lPJaFyH48R)

2. Click on install

![](/files/hPN2So07B89NXtf71GCO)

3. Select the channel, where you want @Dot to respond to user requests and confirm

![](/files/we6Rp49psqNqBclwpNct)

4. Chat with @Dot

<figure><img src="/files/w99QjuKCb258JsntEbGw" alt=""><figcaption></figcaption></figure>

### Slack Data Handling

What Dot reads

* Only messages where it’s invoked (@Dot) and, by default, subsequent messages in that thread while Dot is active there.
* Dot does not read unrelated channel messages.

Private channels

* Adding Dot to a private channel does not grant it access to past history.
* Dot can only see messages where it’s invoked or within threads it’s participating in.
* Non-Dot messages remain invisible to Dot.

Direct messages (optional)

* DMs can be enabled by admins. These interactions are stored like channel interactions and are auditable.

Retention & admin audit

* We securely store only interactions involving Dot to provide context and auditability.
* Admins can review them in Admin Console → History.
* Workspace data deletion is available on request.

## Tips for talking with Dot

1. To start a new topic, begin a new thread with `@Dot` in a channel.
2. Follow-up questions in the same thread use the full thread as context — no need to tag `@Dot` again.
3. If Dot didn't get it right, just rephrase your question.
4. You can upload files (CSVs, spreadsheets) and Dot will analyze them.
5. Ask Dot to create charts, generate PowerPoint presentations, or schedule recurring reports.

### Next Steps

See [Channel Routing](/integrations/slack-and-teams/channel-routing) to route specific channels to different workspaces.


# Reinstall Slack App

How to remove and reinstall the Dot Slack app

Sometimes Dot's Slack app needs to be reinstalled — for example, when we update the app's permissions or fix a configuration issue. This guide walks you through the process step by step.

{% hint style="info" %}
Reinstalling the app does **not** delete your chat history or any data in Dot. It only resets the Slack connection.
{% endhint %}

## Before You Start

**Note which channels Dot is currently in.** After reinstalling, Dot will need to be re-added to each channel manually.

You can find the full list on the Dot app's configuration page in Slack (see Step 1) — it shows all channels under **Bot user**:

<figure><img src="/files/5xJkR0UsgxQPYZBjbsy0" alt="The Slack app configuration page showing which channels Dot is in and the Remove App button"><figcaption><p>The Bot user section lists all channels Dot is currently in. Note these down before removing.</p></figcaption></figure>

{% hint style="warning" %}
Only **Slack Workspace Owners** and users with app management permissions can remove and install apps.
{% endhint %}

## Step 1: Remove the Dot App from Slack

1. In Slack, click your **workspace name** in the top-left corner
2. Go to **Tools & settings** → **Manage apps**
3. Click **Apps** in the left sidebar and find **Dot**

<figure><img src="/files/S4agpCss6BPd5ORNCWFS" alt="Slack Apps page showing Dot in the list"><figcaption><p>Find Dot in your installed apps.</p></figcaption></figure>

4. Click on **Dot**, then click **Configuration**

<figure><img src="/files/OCNbOEcE1ayvMSTqO9Y3" alt="Dot app detail page with Configuration link"><figcaption><p>Click Configuration to open the app settings.</p></figcaption></figure>

5. Scroll down to **Remove app** and click **Remove app**
6. Confirm the removal when prompted

## Step 2: Reinstall Dot via Settings

1. Go to [Dot Settings](https://app.getdot.ai/settings)
2. Click on the **Slack** integration
3. Click **Add Dot to Slack**

<figure><img src="/files/X8J7gFiuWORcTpawMsez" alt="Dot Settings page showing the Slack integration panel"><figcaption><p>Click "Add Dot to Slack" to start the reinstallation.</p></figcaption></figure>

4. In the Slack authorization screen, select the channel where Dot should respond by default and click **Allow**

## Step 3: Re-add Dot to Your Channels

After reinstalling, Dot is only in the default channel you selected during installation. You need to manually re-add Dot to any other channels it was previously in.

To add Dot to a channel:

1. Open the Slack channel where you want Dot
2. Type `/invite @Dot` in the message field and press Enter

Repeat this for every channel from your list in the "Before You Start" step.

{% hint style="info" %}
You can also add Dot to a channel by typing `@Dot` in the channel — Slack will prompt you to invite the app.
{% endhint %}

## Troubleshooting

### Dot doesn't appear in the install screen

Make sure the previous app was fully removed from Slack (Step 1). If the old installation is still active, the reinstall may not work correctly.

### Dot isn't responding in a channel

Verify that Dot has been invited to the channel (Step 3). Dot can only respond in channels where it has been explicitly added.

### Permission errors during install

Contact your Slack Workspace Owner — they may need to approve the Dot app before it can be installed.


# Microsoft Teams

Use Dot where you work with your team

## Cloud Setup

1. In Settings Click on Add Dot to Teams

<figure><img src="/files/siqZPyRtN2XAB0o9rSeH" alt="" width="292"><figcaption></figcaption></figure>

2. Install App in Teams App Store
3. Add Dot to group channel or chat with @Dot EU/US
4. Say `@Dot EU hello`
5. Copy the tenant id to settings and save
6. Chat! 🎉

## Self-hosted Setup

Please contact us at <hi@getdot.ai>.

This setup requires to manually add the Teams app with the correct endpoints to your Tenant.

## Tips for talking with Dot

1. To start a new topic, begin a new conversation with `@Dot` in a group channel or chat.
2. In a group channel, follow-up questions can be asked directly in the thread — Dot uses the whole thread as context. No need to tag `@Dot` again.
3. In a direct chat, Dot uses your previous messages as context. Say `@Dot reset` to start fresh.
4. If Dot didn't get it right, just rephrase. You can also check which data source Dot picked to understand what it can answer.
5. You can upload files (CSVs, spreadsheets) and Dot will analyze them.
6. Ask Dot to create charts, generate PowerPoint presentations, or schedule recurring reports to your email or Slack.

## Permissions

Dot uses group-based permissions to control access. This means a user’s ability to access data via Dot depends on the groups they belong to.

In Microsoft Teams, Dot distinguishes between two types of user contexts:

**1. Direct Messages**

* Dot is used in a 1:1 conversation with an individual user.
* The user is identified by their email address (e.g., <claude.shannon@bell.com>).
* On first interaction, Dot automatically creates a user profile using default settings.<br>

**2. Chats and Team Channels**

* Dot is used in a group context—either a group chat or a team channel.
* Dot creates a shared user profile for the group or channel, not for each individual.
* All users within that group or channel will have the same permissions in Dot.
* This is useful for managing access by organizational domain or function, such as having separate permissions for a marketing channel and a finance channel.
* Example user ID for a team channel:

<pre><code><strong>teams_19:04ca1344-6505-41l3-b743-e64bcf1124f8_5d982b38-eb38-46b6-b929-42ff0392d01a@unq.gbl.spaces@getdot.ai
</strong></code></pre>

Admins can assign different groups (and thus different permissions) to these shared group-level users, allowing tailored access control without managing individual users within each channel.

## Next Steps

See [Channel Routing](/integrations/slack-and-teams/channel-routing) to route specific channels to different workspaces.


# Channel Routing

Route channels and DMs to specific workspaces

Each Slack workspace or Teams tenant connects to one Dot organization or workspace. Use channel routing to direct specific channels to different workspaces with their own data access.

### Two Approaches

* **Route channels within one connection** — Connect Slack/Teams to your main org, then route specific channels to workspaces (e.g., #sales → Sales workspace with CRM access)
* **Separate connections per workspace** — Connect different Slack workspaces or Teams tenants directly to different Dot workspaces

### How Routing Works

| Message type | Routing logic                                                                     |
| ------------ | --------------------------------------------------------------------------------- |
| **DMs**      | Routed by sender's email to their preferred workspace                             |
| **Channels** | Each channel gets a virtual user — assign it to a workspace to route all messages |

### Route a Channel to a Workspace

1. **@mention Dot in the channel** — this automatically creates a channel user (e.g., `#sales-analytics · Slack channel`) in your main organization's user list. You don't need to create anything manually.
2. **Add the channel user to the target workspace** — go to **Settings → Workspaces**, find the target workspace, and click **Manage Users**. In the search dropdown, search for the channel name (e.g., `#sales`) and add it.
3. **Set the channel's default workspace** — go to **Settings → Users**, find the channel user in the list, and change the **Default Workspace** dropdown to the target workspace.
4. **Test** — send another message in the channel and verify Dot responds using the workspace's data.

<figure><img src="/files/8cnzWDRa2OxJmOzhO59H" alt=""><figcaption><p>Set the default workspace for a channel or user</p></figcaption></figure>

{% hint style="info" %}
Don't use the "Add User" button on the Users tab — that's for inviting new users by email. Channel users are created automatically when you @mention Dot. You just need to find them and assign them to a workspace.
{% endhint %}

### Route Individual Users

Same process: add them to the target workspace via **Settings → Workspaces → Manage Users**, then set their default workspace in **Settings → Users**.

### Troubleshooting

| Problem                             | Fix                                              |
| ----------------------------------- | ------------------------------------------------ |
| Messages going to wrong workspace   | Check user/channel's default workspace setting   |
| Bot not responding in channel       | @mention Dot (required in channels, not DMs)     |
| Bot not responding in Teams channel | Verify bot was added to the channel              |
| "Already connected" error           | That Slack/Teams is connected to another Dot org |


# Email

Ask Dot questions straight from your inbox

Dot can answer questions over email, just like in the web app, Slack, or Teams. Send a question to Dot's email address and the answer comes back in the same thread — with charts inline, data as attachments, and a link to keep exploring in the browser.

## Finding Dot's email address

Open **Settings → Connections → Email**. When the email bot is enabled, the card shows **"Email Dot directly at …"** with a copy button. The address is region-specific, so always copy it from there rather than typing it from memory.

{% hint style="info" %}
The card (and the address) only appears once an admin has switched on **Enable Email Bot** in the same place.
{% endhint %}

The first time you email from a new address, Dot creates your account automatically — no verification step needed. Controlling the mailbox is proof enough that the message is from you.

## Asking a question

* **New question** – send an email with your question in the body.
* **Follow-up** – reply to Dot's answer. Your reply stays in the same conversation, so Dot keeps the context of the thread.
* **Forwarding** – forward an email and Dot treats the forwarded content as your question (it is not stripped away as a quote).
* **Attachments** – attach a CSV or file and ask about it; Dot uses it as context.

## What you get back

* A reply **in the same email thread**, with any charts shown inline and data exported as CSV/file attachments.
* A **"View the complete analysis" link** that opens the full, interactive result in the web app.
* For **longer-running questions**, an immediate **acknowledgement email** with a follow-progress link, so you can watch the answer come together in the browser instead of waiting in silence. The finished answer then lands in the same thread. Quick questions skip the acknowledgement and simply get the answer.

{% hint style="info" %}
Each email thread belongs to the person who started it. Replies from other people to that thread aren't processed, and automatic out-of-office replies are ignored.
{% endhint %}


# CSV & Excel Files

Drop a CSV or Excel file into Dot and query it like a table.

Not everything worth analyzing lives in a database. Board exports, vendor reports, one-off extracts, the spreadsheet finance mails around — drop them into Dot and ask your questions, without waiting for anyone to load them into the warehouse first.

There are two ways to use files, depending on how long they should live.

## Attach a file to a chat

Attach a CSV or Excel file directly to your question in the web app. Dot reads it, profiles the columns, and analyzes it in place — ideal for one-off questions where the file *is* the dataset.

## Upload as a lasting table

For files the whole team should query repeatedly, an admin can upload them under **Settings → Connections → Upload Files**. The file becomes a table in your Model like any other source: Dot detects column names and types automatically, you can add descriptions and relationships, and everyone can ask questions against it.

Supported formats: `.csv`, `.xlsx`, and `.xls`.

{% hint style="info" %}
Files are a great on-ramp, not a governance strategy. When a spreadsheet becomes a system of record, move it to a governed source — or connect it as a [Google Sheet](/integrations/databases/google-sheets) so at least everyone reads the same live version.
{% endhint %}


# Single Sign On

No more forgotten passwords


# Azure Active Directory

Single Sign-On and permission mapping with Microsoft Entra ID (Azure AD)

Dot integrates with Microsoft Entra ID (Azure AD) using OAuth 2.0 / OpenID Connect. You can use it for sign-in only, or go further and drive a user's **role, Dot groups, workspace memberships, and access scope** directly from their Azure AD group membership on every login.

This guide is in three parts:

1. **Part 1 - Azure configuration** - register the application in Azure.
2. **Part 2 - Connect Dot to Azure** - enter the credentials in Dot.
3. **Part 3 - Group mappings (optional)** - map Azure AD groups to Dot roles, groups, and workspaces, and control how the rollout reaches your users.

***

## Part 1: Azure Configuration

### Step 1: Register a New Application in Azure

1. Go to the Azure portal and navigate to **Azure Active Directory** > **App registrations**.
2. Click on **New registration**.

### Step 2: Application Registration

1. Enter the name of the application, for example, `Dot Azure SSO`.
2. Under **Supported account types**, select the relevant option for your organization (single tenant is the typical choice).
3. For the **Redirect URI**, select **Web** and enter the URI shown in Dot's **Azure Entra ID** card (Settings > Connections), where it appears as a read-only **Redirect URI** field. This URI is unique to your organization, so copy it exactly as Dot displays it.

<figure><img src="/files/9qlLBY5ey0XTIZ7C8m9A" alt=""><figcaption></figcaption></figure>

### Step 3: Application Overview

1. Once the application is registered, you will be redirected to the application's overview page.
2. Copy the **Application (client) ID** and **Directory (tenant) ID** and save them for later use.

<figure><img src="/files/NIPbD0Ql0cUuj7lgiXSo" alt=""><figcaption></figcaption></figure>

### Step 4: Certificates & Secrets

1. In the application's menu, click on **Certificates & secrets**.
2. Click on **New client secret**.
3. Add a description for the secret and set an expiry as required (12 or 24 months is recommended).
4. Once created, **immediately copy the secret Value** (not the Secret ID) - it is only shown once.

<figure><img src="/files/gUrWsjXBUTzCCWzlk5dJ" alt=""><figcaption></figcaption></figure>

### Step 5: API Permissions

If you only need sign-in, the default `openid` / `email` / `profile` scopes are enough and you can skip ahead.

To use **group mappings** (Part 3), Dot reads the signed-in user's own group membership through Microsoft Graph. Add the delegated permission below:

| Permission  | Type      | Purpose                                                                     | Admin consent        |
| ----------- | --------- | --------------------------------------------------------------------------- | -------------------- |
| `User.Read` | Delegated | Read the signed-in user's profile and check their group membership on login | Usually not required |

{% hint style="info" %}
On strict tenants, users may see a consent prompt they are not allowed to accept. In that case a tenant admin grants consent once in **Azure Portal > Enterprise Applications > \[Dot] > Permissions > "Grant admin consent"**, and users will no longer be prompted.
{% endhint %}

### Step 6: Restrict Who Can Sign In (Optional)

To limit which users can use Dot SSO rather than allowing everyone in your tenant:

1. In Azure, go to **Enterprise Applications** and open the application created for Dot.
2. Go to **Manage** > **Users and groups** and click **Add user/group** to assign the users or groups that should have access.

<figure><img src="/files/nuavyUprhOESy4SF66YB" alt=""><figcaption></figcaption></figure>

3. Go to **Manage** > **Properties**.

<figure><img src="/files/s4LoRn5HFnRwMFkj66Qe" alt=""><figcaption></figcaption></figure>

4. Set **Assignment required?** to **Yes**. Now only explicitly assigned users or groups can authenticate.

***

## Part 2: Connect Dot to Azure

1. Log into Dot as an **admin** and open **Settings > Connections**.
2. Open the **Azure Entra ID** card and fill in:
   * **Client ID** - the Application (client) ID from Step 3
   * **Client Secret** - the secret value from Step 4 (leave blank later to keep the existing secret)
   * **Metadata URL**, built from your tenant ID: `https://login.microsoftonline.com/{tenant-id}/v2.0/.well-known/openid-configuration`
3. Click **Save** and test the sign-in.

Once SSO is active, the **Group mappings** section appears in the same card.

***

## Part 3: Group Mappings

Group mappings let Azure AD decide what a user can do in Dot. Without any mappings, SSO is sign-in only and permissions stay exactly as they are managed inside Dot today.

<figure><img src="/files/g2pOKP3MIdNFmQOxfA48" alt="The Azure Entra ID card with SSO credentials, group mappings, rollout controls, and the mappings table"><figcaption><p>The Azure Entra ID card: SSO credentials at the top, then the group-mapping controls</p></figcaption></figure>

### How It Works

1. On every SSO login, Dot checks the user's Azure AD group membership against your configured mappings.
2. All matching mappings are combined into one result:
   * **Role** - the highest-privilege role across all matched groups wins (Admin > Modeler > User).
   * **Dot groups** - the union of the Dot groups from every matched mapping.
   * **Workspaces** - the union of workspace memberships; if the same workspace appears in more than one matched group, the highest role for that workspace wins.
   * **Access scope** - if any matched group restricts the user to **Workspaces only**, that restriction applies; otherwise the user gets **Full** access.
3. If a managed user matches no mapping, they receive the **fallback role**.
4. Membership is re-evaluated on every login, so Azure AD changes take effect the next time the user signs in.

### Dot Roles

| Role        | Typical permissions                                 |
| ----------- | --------------------------------------------------- |
| **Admin**   | Full access, including settings and user management |
| **Modeler** | Curate and model data for a workspace               |
| **User**    | Ask questions and consume answers                   |

### Adding a Mapping

In the **Azure Entra ID** card, under **Mappings**, click **Add mapping** and fill in:

| Field                  | Meaning                                                                                                                                                     |
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Azure group ID**     | The Object ID (UUID) of the Azure AD security group. Find it in **Azure AD > Groups > \[group] > Overview**.                                                |
| **Group display name** | Optional friendly label for your own reference.                                                                                                             |
| **Role**               | Org-level role to grant (Admin, Modeler, User), or none if this group only controls workspaces.                                                             |
| **Dot groups**         | Comma-separated Dot group names to add the user to (for example `marketing, exec`).                                                                         |
| **Workspaces**         | One or more workspaces, each with a role, that the user should belong to.                                                                                   |
| **Access scope**       | **Inherit** (do not impose a scope from this group), **Full** (whole workspace/org), or **Workspaces only** (restrict the user to their mapped workspaces). |

<figure><img src="/files/5osKBLUdCzv7P3dWj338" alt="The Add mapping dialog with fields for Azure group ID, role, Dot groups, workspaces, and access scope"><figcaption><p>Adding a mapping</p></figcaption></figure>

#### Example

| Azure AD Group | Role  | Dot groups | Workspaces   | Access scope    |
| -------------- | ----- | ---------- | ------------ | --------------- |
| Dot Admins     | Admin | -          | -            | Full            |
| Sales Analysts | User  | sales      | Sales (User) | Workspaces only |

A member of "Dot Admins" becomes an Admin with full access. A member of "Sales Analysts" becomes a User in the Sales workspace, added to the `sales` Dot group, and limited to that workspace. A user in both groups becomes an Admin (highest role wins).

<figure><img src="/files/Ac3a3yMzvtL6XYfFZ9GK" alt="The mappings table showing two configured mappings with their roles, Dot groups, workspaces, and access scope"><figcaption><p>Configured mappings</p></figcaption></figure>

### Fallback Role

The **Fallback role** (User, Modeler, or Admin) is applied to managed users who sign in but match none of your mappings. Set this to the least-privileged role that still makes sense for your organization.

### Workspace Membership: Add Only vs. Add and Remove

The **Workspace membership** setting controls how aggressively Dot syncs workspace memberships:

| Mode                   | Behavior                                                                                                                                                                    |
| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Add only** (default) | Add users to workspaces when a mapping matches. Memberships are never removed automatically.                                                                                |
| **Add and remove**     | Make Dot match Azure exactly. A user who loses a group mapping loses the corresponding workspace access on their next login. Their per-workspace chat history is preserved. |

Start with **Add only** while you validate your mappings, and switch to **Add and remove** once you want Azure AD to be the single source of truth.

### Restrict to Tenants (Optional)

Under **Restrict to tenants**, enter one Azure tenant ID (UUID) per line to accept logins only from those tenants. Leave it empty to accept any tenant the SSO app trusts.

***

## Controlling the Rollout: Who Is Managed

Turning on group mappings does not have to flip everyone at once. The **Who is managed** controls let you decide, separately, whether mappings apply to new users, existing users, or only a subset of existing users. This is the safest way to introduce Azure-driven permissions without disturbing the people already working in Dot.

There are two independent switches:

| Switch                                                 | Default | Effect                                                                                                                         |
| ------------------------------------------------------ | ------- | ------------------------------------------------------------------------------------------------------------------------------ |
| **Sync new users' permissions from Azure groups**      | On      | Any user who signs in for the first time via SSO has their role, groups, workspaces, and access scope decided by the mappings. |
| **Sync existing users' permissions from Azure groups** | Off     | Users who already exist in Dot keep their current, hand-curated permissions until you opt them in.                             |

When you turn on syncing for existing users, an additional **workspace scope** appears:

> *Only manage existing users who are members of these workspaces. Leave all unchecked to manage every existing user.*

This is the key to a controlled, gradual rollout.

<figure><img src="/files/j498aYgK5uujEwfkLp4M" alt="Who is managed panel with both sync options enabled and the existing-user scope limited to a single checked workspace"><figcaption><p>Existing users managed only in the checked workspace; everyone else keeps the status quo</p></figcaption></figure>

### The Recommended Rollout Pattern: Pilot in One Workspace First

A common and recommended way to adopt group mappings is to pilot them in a single workspace before rolling out organization-wide. The goal is:

* **New users** are managed by Azure from day one.
* **Existing users in the pilot workspace** are managed by Azure, so you can verify the mappings against real accounts.
* **Everyone else, including the main workspace, keeps the status quo** and is untouched until you are ready.

To set this up:

1. In Azure AD, create the security groups for the pilot workspace (for example one group per role) and add the pilot users.
2. In Dot's Entra ID card, add the corresponding **mappings** for those groups, pointing them at the pilot workspace with the appropriate roles.
3. Under **Who is managed**:
   * Keep **Sync new users' permissions from Azure groups** enabled.
   * Enable **Sync existing users' permissions from Azure groups**.
   * In the workspace scope list, check **only the pilot workspace**. Leave all other workspaces, including the main workspace, unchecked.
4. Set the **Fallback role** to a safe, low-privilege role so any managed pilot user who matches no mapping lands somewhere harmless.
5. Keep **Workspace membership** on **Add only** during the pilot.

With this configuration, only existing members of the pilot workspace are evaluated against Azure on login. Existing users everywhere else continue with exactly the permissions they have today. When the pilot is successful, widen the rollout by checking additional workspaces, or clear the workspace scope entirely to manage every existing user.

You can revisit this choice at any time through **Change who is managed**:

<figure><img src="/files/pTs49GkteuRqD09hHF0x" alt="The Who should Azure manage dialog offering Everyone or New users only"><figcaption><p>The rollout chooser, reachable any time via "Change who is managed"</p></figcaption></figure>

{% hint style="warning" %}
The workspace scope only constrains **existing** users. New users are governed by the "Sync new users" switch regardless of which workspaces are checked. Keep that in mind if you want new users held back as well during a pilot.
{% endhint %}

{% hint style="info" %}
Switching from a scoped rollout to managing **Everyone** clears the workspace scope, so from then on all existing SSO users are evaluated against your mappings on their next login. Re-check specific workspaces if you want to narrow it again.
{% endhint %}

***

## Testing

1. Log out of Dot and choose **Sign in with Microsoft** on the login page.
2. Authenticate with an Azure AD account (MFA and Conditional Access policies are enforced by Azure during this step).
3. After sign-in, confirm the user's role, Dot groups, and workspace memberships match what their Azure groups should grant.
4. Verify a few cases:
   * A user in a mapped Admin group has Admin access.
   * A user in a mapped workspace group lands in the right workspace with the right role.
   * A managed user in no mapped group receives the fallback role.
   * An existing user **outside** your rollout scope is unchanged.

***

## Troubleshooting

**"Failed to exchange code for token"** - The client secret is incorrect or expired. Verify it in Dot matches Azure, and create a new secret if it has expired.

**Roles or workspaces are not applied** - Confirm group mappings exist for the user's Azure groups, the group Object IDs match exactly, and the user is in scope under "Who is managed". Have the user sign out and back in, since mappings are evaluated on login.

**A rollout change did not take effect for some users** - Permissions are re-evaluated on the next login, not retroactively. Ask affected users to sign out and back in.

**Redirect URI mismatch** - The Redirect URI in Azure must match the one shown in Dot's **Azure Entra ID** card exactly, including protocol and path.

**Consent prompt the user cannot accept** - A tenant admin needs to grant admin consent once for the application's permissions (see Step 5).

***

## FAQ

**What happens if a user is in multiple mapped groups?**\
The results are merged: the highest-privilege role wins, Dot groups and workspaces are unioned, and any group that restricts access scope to "Workspaces only" applies.

**What if no group mappings are configured?**\
SSO is sign-in only. Permissions stay exactly as managed inside Dot.

**How quickly do Azure AD changes take effect?**\
On the user's next login. Dot reads group membership fresh each time. Azure AD itself may take a few minutes to propagate group changes.

**Can I try mappings on a few people before rolling out to everyone?**\
Yes. Manage new users only, or manage existing users scoped to specific workspaces. See "Controlling the Rollout" above.

**Does it work with MFA and Conditional Access?**\
Yes. Any MFA or Conditional Access policy in your tenant is enforced during the Microsoft login step. No special configuration is needed in Dot.


# Okta

Single Sign On - no more forgotten passwords

## Integrating Single Sign-On (SSO) with Okta for Dot

This guide walks you through the process of creating an app integration in Okta and setting up SSO with Dot.

You can use Okta for sign-in only, or go further and let your **Okta groups become Dot groups**, so group membership is managed in Okta and never maintained twice — see [Group Sync](#group-sync-optional) below.

### Step 1: Create a New App Integration in Okta

1. Log in to your Okta admin dashboard.
2. Navigate to **Applications** > **Applications**.
3. Click on **Create App Integration**.

### Step 2: Configure the App Integration

1. Select the **OIDC - OpenID Connect** option.
2. Choose **Web Application** as the application type and click **Next**.

<figure><img src="/files/Yd6pfrous5XbHNeA8wo2" alt=""><figcaption></figcaption></figure>

### Step 3: Set Redirect URI

1. In a separate browser tab, go to your **Dot Settings** > **Okta** section to copy the Redirect URI.
2. Return to the Okta tab and paste the copied URI into the **Sign-in redirect URIs** field.

<figure><img src="/files/LFue0qrVsfJ7Kl1Sefig" alt=""><figcaption></figcaption></figure>

### Step 4: Configure General Settings

1. Provide a name for your integration, e.g., `Dot SSO Integration`.
2. Add the Redirect URI from the Dot settings to **Sign-in redirect URIs**.
3. Set the **Sign-out redirect URI** to the base domain of Dot:
   * For EU: `https://eu.getdot.ai`
   * For US: `https://app.getdot.ai`
4. Set the logo

{% file src="/files/Yowbt1xpn7sRmnmMHF0z" %}

<figure><img src="/files/c9J7Z57tbXyFWcj554KZ" alt=""><figcaption></figcaption></figure>

### Step 5: Assign Users or Groups

1. Choose to either **Assign** individual users or **Assign to groups** within your organization.

### Step 6: Copy Client Credentials

1. After saving the new app integration, navigate to the **General** tab of your newly created app. 2. Copy the **Client ID** and **Client Secret**.

<figure><img src="/files/mMfnK74xZpfRHHHyOlOV" alt=""><figcaption></figcaption></figure>

### Step 7: Configure Dot with Okta Credentials

1. Go back to your Dot Settings > Okta section.
2. Paste the **Client ID** and **Client Secret** into the respective fields.

### Step 8: Metadata URL

1. The Metadata URL is essential for SSO operations. Construct it using your Okta domain:
   * Format: `https://{okta-url}/.well-known/openid-configuration`
   * Example: `https://dev-12345678.okta.com/.well-known/openid-configuration`
2. Your Okta URL can be found in the drop-down menu under your username at the top right corner of the Okta dashboard.

<figure><img src="/files/G1HvKaf2F15cJ63Tw03a" alt=""><figcaption></figcaption></figure>

### Step 9: Initiate Login URI (Recommended)

By default, clicking the Dot tile on the Okta dashboard lands users on Dot's generic login page, where they still have to type their email address before the SSO button appears. Setting an **Initiate login URI** sends them straight into your organization's SSO flow instead.

1. In Okta, open your Dot app and go to the **General** tab.
2. Click **Edit** on **General Settings**.
3. Set **Initiate login URI** to your organization's direct login URL:
   * `https://eu.getdot.ai/login/{your-org-id}` (EU)
   * `https://app.getdot.ai/login/{your-org-id}` (US)
4. **Save**.

Your organization ID is the value shown in Dot under **Settings** > **Users** (usually your email domain, for example `acme.com`). So a full URI looks like `https://eu.getdot.ai/login/acme.com`.

{% hint style="info" %}
The query form `https://eu.getdot.ai/login?org_id=acme.com` behaves identically, if you prefer it or already have it in circulation. Either URL is also a good bookmark to share with your users directly.
{% endhint %}

{% hint style="success" %}
If your organization has exactly one SSO provider active, landing on this URL takes users all the way to Okta without a further click. With password login still enabled, users can fall back to it from the standard `/login` page.
{% endhint %}

### Finalizing the Integration

After you have entered all the necessary information into Dot's Okta settings:

1. Click **Save** to apply the settings.
2. Test the SSO integration to ensure it's working as expected.

By following these steps, you will have successfully set up SSO with Okta for your Dot application. Ensure that all copied values are kept secure and are only shared with authorized personnel within your organization.

## Group Sync (Optional)

Group sync makes your **Okta groups** the groups Dot uses. A user signing in through Okta is placed in the Dot groups matching their Okta groups, so you manage membership in Okta only.

Unlike the Azure AD and Google integrations, there is **no mapping table to fill in**. Okta group names are used as Dot group names directly, so there is no list to keep in step on the Dot side. What Dot receives is decided in Okta, by a groups claim you add to the Dot app.

Group sync affects **group membership only**. It does not set roles, does not add or remove workspace memberships, and never blocks a login.

### Step 1: Create the Groups in Okta

Group sync mirrors groups that already exist in Okta — it never creates them. Most Okta directories are organized around IT concerns (`vpn-users`, `office-oslo`) rather than data access, so start by making a small set of groups for the access you want in Dot:

1. In Okta, go to **Directory** > **Groups** > **Add group**.
2. Name it with a shared prefix, for example `dot-commercial`.
3. Open the group, click **Assign people**, and add the users who belong to it.
4. Repeat for each group you need (`dot-finance`, `dot-analysts`, …).

The shared prefix is what makes the filter in [Step 3](#step-3-choose-which-groups-dot-receives) simple, and it makes clear at a glance which Dot groups are owned by Okta.

{% hint style="warning" %}
**The prefix comes across with the name.** Group names arrive verbatim, so the Okta group `dot-commercial` becomes the Dot group `dot-commercial` — the prefix is *not* stripped. Scope your tables and explores to the prefixed names, and if you want the Dot group to read exactly `commercial` instead, name the Okta group `commercial` and use a matcher that still selects it.
{% endhint %}

### Step 2: Add a Groups Claim in Okta

Dot reads group membership from the **ID token**, so Okta has to include it there.

1. In the Okta admin dashboard, go to **Applications** > **Applications** and open your Dot app.
2. Open the **Sign On** tab.
3. Scroll to **Token claims** and expand **Show legacy configuration**.
4. Next to **Group Claims**, click **Edit**.
5. Leave **Groups claim type** as **Filter**.
6. Under **Groups claim filter**, keep the claim name `groups`, then pick a matcher and enter a value that selects the groups Dot should see (see [Step 3](#step-3-choose-which-groups-dot-receives)).
7. **Save**.

<figure><img src="/files/ct8ZUlKjqwRCB7DGl8fO" alt="The Group Claims form in Okta with claim type Filter, claim name groups, and a Starts with dot- filter"><figcaption><p>Group Claims under <strong>Show legacy configuration</strong>: claim name <code>groups</code>, filtered to groups starting with <code>dot-</code></p></figcaption></figure>

{% hint style="warning" %}
**The filter needs a value.** The claim name defaults to `groups` and the matcher to **Starts with**, but the value box starts empty. An empty value displays as **Groups claim filter: None** and sends no groups at all — the most common reason group sync appears to do nothing.
{% endhint %}

{% hint style="info" %}
**Can't find Group Claims?** In current Okta versions the group-claim fields are not on the **OpenID Connect ID Token** card (that card only holds Issuer and Audience). They live under **Token claims** > **Show legacy configuration**. The newer expression-based **Token claims** editor above it is not needed for group sync.
{% endhint %}

{% hint style="warning" %}
The claim must be on the **ID token**. A claim added only to the access token or only to the `/userinfo` endpoint will not reach Dot, and group sync will behave as if the user is in no groups.
{% endhint %}

{% hint style="info" %}
The steps above apply to the **Okta org authorization server**, which is what the Metadata URL in [Step 8](#step-8-metadata-url) points at (`https://{okta-url}/.well-known/openid-configuration`). If you pointed Dot at a **custom authorization server** instead (a Metadata URL containing `/oauth2/{id}/`), the Sign On tab has no effect — add the `groups` claim under **Security** > **API** > **Authorization Servers** > *\[your server]* > **Claims**, with **Include in token type** set to **ID Token**.
{% endhint %}

### Step 3: Choose Which Groups Dot Receives

The filter is your control over what Dot gets. Keep it narrow — send only the groups that should drive access in Dot.

| Matcher           | Value      | Sends                                                       |
| ----------------- | ---------- | ----------------------------------------------------------- |
| **Starts with**   | `dot-`     | Only groups whose name begins with `dot-`. **Recommended.** |
| **Equals**        | `analysts` | That one group.                                             |
| **Matches regex** | `.*`       | Every group the user belongs to.                            |

{% hint style="info" %}
`Matches regex` `.*` works, but sends every group in your directory that the user belongs to. In a large directory that makes the token big and fills Dot with groups that mean nothing there. A prefix like `dot-` keeps the set deliberate.
{% endhint %}

{% hint style="success" %}
The claim filter and the app assignment do different jobs. Okta only authenticates people the Dot app is assigned to (**Applications** > **Dot** > **Assignments**), so that governs **who can sign in**. The groups claim governs **what they can see** once inside. Widening the filter never grants anyone a login.
{% endhint %}

### Step 4: Turn On Group Sync in Dot

1. In Dot, go to **Settings** > **Connections** and open the **Okta** card.
2. Under **Group sync**, switch **Use Okta groups as Dot groups** on.

The setting only appears once Okta SSO is configured and saved.

{% hint style="info" %}
With group sync on, Dot requests one extra scope at login (`groups`), which is what makes Okta emit the claim. The filter from Step 2 shapes *which* groups it returns, but without the scope Okta sends none at all — so both halves are required. Dot handles the scope automatically; there is nothing to configure in Okta for it. Users may see a one-time consent prompt. Turning the toggle back off drops the scope again.
{% endhint %}

### Step 5: Check It Worked

Nothing is synced until a user signs in again, so verify before scoping any data to the new groups:

1. Sign out of Dot completely, then sign back in through Okta.
2. Go to **Settings** > **Users**.
3. The synced groups appear against each user, alongside any groups assigned by hand.

If they are missing, work through it in this order:

1. Is **Group sync** on in Dot's Okta card? It also controls the `groups` scope, so with it off Okta sends no groups no matter how the filter is set.
2. Is **Groups claim filter** showing a value rather than **None**? (Okta > **Sign On** > **Token claims** > **Show legacy configuration**.)
3. Does the group name actually match your claim filter? A group excluded by the filter never reaches Dot, even though the user is in it.
4. Was it a full sign-out and sign-in? An existing session is not re-evaluated.

### Step 6: Use the Groups

Synced groups behave exactly like groups created in Dot, so you can scope data with them. To restrict a table or Looker explore to a group:

1. Go to **Model** and click the table or explore.
2. Open the **Access** tab.
3. Add the groups that should have access, and remove `all_users` if it should no longer be visible to everyone.

Users then only see the tables and explores their groups grant. New tables default to `all_users`, so they are visible to everyone until scoped.

### How the Sync Behaves

* **Applied at every login.** Changes in Okta take effect the next time the user signs in, not immediately.
* **Names are used as-is**, lowercased and prefix included. An Okta group `Dot-Commercial` becomes the Dot group `dot-commercial`. Matching ignores case.
* **Removing someone from an Okta group** removes the matching Dot group on their next sign-in.
* **Groups you assign by hand in Dot are left alone** — sync only manages the groups it added. The exception is a name that is both hand-assigned and sent by Okta: Okta owns it, so removing it in Okta removes it in Dot.
* **Workspace identities are kept in step too.** A user who is a member of a workspace has their groups synced there as well, so revoking an Okta group also revokes the workspace access it granted. Roles and workspace memberships themselves are never changed.

### Troubleshooting

**Nobody gets any groups** — Almost always a missing or misplaced claim. In Okta, open **Sign On** > **Token claims** > **Show legacy configuration** and check **Group Claims**. If **Groups claim filter** reads **None**, the filter has no value and sends nothing. Also confirm the claim is named `groups`, is on the **ID token** (not the access token or `/userinfo`), and that the filter actually matches your group names. Have the user sign out and back in afterwards.

**A group is missing for one user** — Check they are a member of that group in Okta, and that the group matches your claim filter. A group excluded by the filter never reaches Dot, even though the user is in it.

**Groups arrived but the user still cannot see a table** — Group sync grants group membership, not data access. Scope the table or explore to that group under **Model** > *\[table]* > **Access**.

**Users lost their groups unexpectedly** — If the claim is removed or renamed in Okta, Dot can no longer tell "in no groups" from "not configured". It removes the groups it had synced rather than leaving access standing on information it can no longer confirm, and records an error in the logs. Logins are not blocked. Restoring the claim restores the groups on the next sign-in.

{% hint style="info" %}
Group sync does not deprovision accounts. Removing a user from the Dot app in Okta stops them signing in, but their Dot account remains until an administrator deletes it in **Settings** > **Users**.
{% endhint %}


# Google

Single Sign-On and permission mapping with Google Workspace

Dot integrates with Google using OAuth 2.0 / OpenID Connect. You can use it for sign-in only, or go further and drive a user's **role, Dot groups, and workspace memberships** directly from their Google group membership on every login — see [Group Sync](#group-sync-optional) below.

## Integrating Single Sign-On (SSO) with Google for Dot

This guide walks you through setting up Google as an SSO provider for Dot using Google Cloud's OAuth 2.0 credentials.

### Step 1: Open Google Cloud Console

1. Go to the [Google Cloud Console](https://console.cloud.google.com/).
2. Select an existing project or create a new one for your organization.

### Step 2: Configure the OAuth Consent Screen

1. Navigate to **APIs & Services** > **OAuth consent screen**.
2. Select **Internal** (recommended for Google Workspace organizations — only users in your organization can sign in) or **External**.
3. Fill in the required fields: App name, User support email, and Developer contact email.
4. Click **Save and Continue** through the remaining steps.

### Step 3: Create OAuth 2.0 Credentials

1. Navigate to **APIs & Services** > **Credentials**.
2. Click **Create Credentials** > **OAuth client ID**.
3. Select **Web application** as the application type.
4. Give it a name, e.g., `Dot SSO`.

### Step 4: Set the Redirect URI

1. In a separate browser tab, go to your **Dot Settings** and click on the **Google** card under Authentication.
2. Copy the **Redirect URI** shown at the top of the form.
3. Return to Google Cloud Console and paste it into the **Authorized redirect URIs** field.

<figure><img src="/files/tlG1cDsqvFGBjkL9omP3" alt=""><figcaption><p>The Dot Google SSO configuration form showing the Redirect URI to copy</p></figcaption></figure>

### Step 5: Copy Client ID and Client Secret

1. After creating the OAuth client, Google will display the **Client ID** and **Client Secret**.
2. Copy both values — you'll need them in the next step.

### Step 6: Configure Dot with Google Credentials

1. Go back to your **Dot Settings** > **Google** section.
2. Paste the **Client ID** and **Client Secret** into the respective fields.
3. Click **Save** to apply the settings.

{% hint style="info" %}
The Google metadata URL (`https://accounts.google.com/.well-known/openid-configuration`) is automatically configured by Dot — no manual entry needed.
{% endhint %}

### Finalizing the Integration

Test the SSO integration by signing out and signing back in with Google.

By following these steps, you will have successfully set up SSO with Google for your Dot application. Ensure that all copied values are kept secure and are only shared with authorized personnel within your organization.

## Optional: Restrict Which Users Can Sign In

If you selected **Internal** for the OAuth consent screen in Step 2, only users within your Google Workspace organization can sign in — no further restriction is needed.

If you selected **External**, you can restrict access by:

1. Going to **APIs & Services** > **OAuth consent screen** in Google Cloud Console.
2. Under **Test users**, adding only the specific email addresses that should have access (while the app is in "Testing" status).
3. To allow all users, submit the app for verification to move it to "Production" status.

## Group Sync (Optional)

Group mappings let Google Workspace decide what a user can do in Dot. Without any mappings, SSO is sign-in only and permissions stay exactly as they are managed inside Dot today.

On every SSO login, Dot resolves the signing-in user's Google group membership (including **nested groups**) and applies the grants of every mapped group they belong to. Membership is re-evaluated on each login, so changes in Google Admin console take effect the next time the user signs in.

<figure><img src="/files/T9YqdZRKN586mq5lSYJO" alt="The Google card with SSO credentials, the group mappings section, and rollout controls"><figcaption><p>The Google card: SSO credentials at the top, then the group-mapping controls</p></figcaption></figure>

### Google-Side Requirements

Two things must be true in your Google environment before enabling any rollout:

1. **Enable the Cloud Identity API** on the Google Cloud project that owns your OAuth client (**APIs & Services** > **Library** > *Cloud Identity API* > **Enable**). Dot resolves group membership through this API at login; if it is disabled, sign-in fails for managed users.
2. **Mapped groups must be visible to their members.** Dot queries membership with the signing-in user's own token, so in **Google Admin console** > **Directory** > **Groups** > *\[group]* > **Access settings**, allow group members to view the member list (this is the default for most groups).

{% hint style="info" %}
Group sync works on **all Google Workspace editions**. On Enterprise and Cloud Identity Premium editions Dot uses Google's transitive membership API; on Business-tier editions it automatically falls back to resolving nested groups step by step. No configuration is needed either way.
{% endhint %}

When group sync is rolled out, Dot requests one additional read-only OAuth scope at login (`cloud-identity.groups.readonly`). Users may see a one-time consent prompt for it.

### How Grants Work

A mapping binds one Google group (by its **email address**) to a list of **grants**. Each grant gives that group a role — and optionally Dot groups — in one workspace:

* Grant the **Main workspace** to give access to the main organization. Without a Main grant, members have no main-org access.
* Grant other workspaces to add members there with the given role.

When a user matches several mappings, the results merge: the **highest role wins** per workspace (Admin > Modeler > User) and Dot groups are unioned.

<figure><img src="/files/euIP4P8FINF8ut4OJBQm" alt="The Add mapping dialog with the Google group email, display name, and a grant targeting the Main workspace with the Admin role"><figcaption><p>Adding a mapping: group email + grants</p></figcaption></figure>

| Field                  | Meaning                                                                                |
| ---------------------- | -------------------------------------------------------------------------------------- |
| **Google group email** | The group's email address, from **Google Admin console** > **Directory** > **Groups**. |
| **Group display name** | Optional friendly label for your own reference.                                        |
| **Grants**             | One or more rows of workspace + role + optional Dot groups.                            |

### Controlling the Rollout: What Google Manages

Mappings are inactive until you choose a rollout scope:

| Scope                                | Effect                                                                                                                                                                                               |
| ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Off** (default)                    | Sign-in only. Mappings are configured but not applied.                                                                                                                                               |
| **Only specific workspaces (pilot)** | Mappings apply only inside the selected workspaces. The main workspace and all other workspaces stay untouched; users in no mapped group are left unmanaged. This is the recommended starting point. |
| **All of Dot**                       | Google becomes the source of truth for the main workspace and every workspace. **Anyone not in a mapped group is denied access.**                                                                    |

{% hint style="warning" %}
**All of Dot** is a default-deny mode. Dot refuses to enable it until at least one mapping grants **Admin on the Main workspace**, so you cannot lock out every administrator. Validate your mappings in a pilot first.
{% endhint %}

### Workspace Membership: Add Only vs. Add and Remove

Under **Advanced**, the **Workspace membership** setting controls how aggressively Dot syncs workspace memberships:

| Mode                   | Behavior                                                                                                                                                                 |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Add only** (default) | Add users to workspaces when a mapping matches. Memberships are never removed automatically.                                                                             |
| **Add and remove**     | Make Dot match Google exactly. A user who loses their mapping loses the corresponding workspace access on the next login. Their per-workspace chat history is preserved. |

### Exempting Individual Users

To take one user out of Google governance without changing any mapping, open **Settings** > **Users**, click the **⋯** menu on the user, and choose **Stop Google-managing this user**. Their permissions are then managed inside Dot again until you re-enable it.

### Troubleshooting

**"Couldn't verify your Google Workspace group membership"** — Dot could not resolve the user's groups, and denies the login rather than guessing (fail-closed). Check that the Cloud Identity API is enabled on the OAuth client's project and that the mapped groups let members view the member list.

**Roles or workspaces are not applied** — Confirm the mapping's group email matches exactly, the user is (transitively) a member, and a rollout scope is selected. Have the user sign out and back in, since mappings are evaluated on login.

**A user was denied after enabling All of Dot** — In full rollout, anyone in no mapped group is denied by design. Add the user to a mapped Google group, or switch back to a pilot scope.


# Generic OIDC

Single Sign On with any OpenID Connect provider

## When to use this provider

Dot ships with dedicated tiles for the SSO providers most customers use — **Azure Active Directory**, **Okta**, and **Google**. If your identity provider isn't one of those — e.g. an in-house IdP, Auth0, Keycloak, Ping Identity, OneLogin, or JumpCloud — use the **Generic OIDC** tile instead. It speaks the same standard — OpenID Connect — and works with any IdP that publishes a discovery document.

If you're using one of the named providers, follow the dedicated guide instead — those use the same plumbing under the hood, but with the IdP-specific URL prefilled and a fixed button label.

## What you'll need from your IdP

Before opening the Dot admin panel, have these three values ready from your IdP:

* **Client ID** — issued when you register Dot as an OAuth/OIDC application
* **Client Secret** — issued alongside the Client ID; treat it like a password
* **Metadata URL** — also called the **OpenID configuration URL** or **discovery URL**. Always ends in `/.well-known/openid-configuration`. Examples:
  * Auth0: `https://{your-tenant}.auth0.com/.well-known/openid-configuration`
  * Keycloak: `https://{host}/realms/{realm}/.well-known/openid-configuration`
  * In-house / other: ask your IdP team for the discovery URL

The metadata URL is what makes this generic — Dot fetches the IdP's signing keys, authorization endpoint, and token endpoint from that one URL, so you don't have to enter them by hand.

## Step 1: Open the Generic OIDC tile in Dot

1. Sign in to Dot as an admin.
2. Open **Settings** → **Connections** → **Authentication**.
3. Click the **Generic OIDC** card.

<figure><img src="/files/pX0BbRw89ymw4JsWCiEW" alt=""><figcaption><p>The Generic OIDC tile in Dot's admin Connections tab</p></figcaption></figure>

## Step 2: Copy Dot's Redirect URI into your IdP

The first field in the form is a read-only **Redirect URI** that looks like:

```
https://app.getdot.ai/api/auth/{your-org-id}/oidc/auth
```

In your IdP, register a new OAuth/OIDC application (sometimes called a "client" or "relying party") and paste this exact URL into the **Redirect URI** / **Callback URL** field. Most IdPs require this to match character-for-character — including the trailing path — or the login will fail with a redirect mismatch error.

If your IdP asks for the OAuth grant type, choose **Authorization Code** (with PKCE if offered). Dot doesn't use implicit or password grants.

## Step 3: Fill in the form in Dot

Back in the Dot admin form, fill in the remaining fields:

* **Client ID** — paste from your IdP
* **Client Secret** — paste from your IdP
* **Metadata URL** — the `/.well-known/openid-configuration` URL
* **Button Label** *(optional)* — what gets displayed on the login page. For example, set this to `Acme SSO` to render a button labelled **Sign in with Acme SSO**. Leave empty to fall back to the generic label **Sign in with SSO**.

<figure><img src="/files/mIJCbB4ufDcznl9gG1Ln" alt=""><figcaption><p>Filled-in Generic OIDC tile, with a custom button label set to <code>Acme SSO</code></p></figcaption></figure>

Click **Save**. After saving, the tile flips to **Connected**, the secret is masked, and the **Remove** button appears next to **Save** so you can disable the integration later.

<figure><img src="/files/XLjxpJqjrkFT9UybLkoF" alt=""><figcaption><p>Saved Generic OIDC tile — secret is masked and the tile is marked Connected</p></figcaption></figure>

## Step 4: Verify the login button

Sign out of Dot, then open the login page for your organization. You should see your **Sign in with {Button Label}** button next to the email/password form.

<figure><img src="/files/yi93gfbsT034zjBpGA5v" alt=""><figcaption><p>The login page renders the Generic OIDC button with the custom label</p></figcaption></figure>

Click it once to verify the round-trip: Dot redirects you to your IdP, you authenticate, the IdP redirects back to Dot, and you land in the app signed in.

## How user provisioning works

The first time someone signs in through Generic OIDC, Dot reads the `email` claim from the ID token and creates an account for them if one doesn't exist yet. After that, sign-ins from the same email reuse the same account.

What role a new person gets depends on one org setting. By default Dot makes new people admins. If you don't want that, turn off "Default new users to admin role" on the Users page before you invite anyone. New people then join as regular users instead. Dot has three roles, Admin, Modeler, and User, and you can change anyone's role later on the Users page. There is no Viewer or Editor role.

{% hint style="warning" %}
With SSO, anyone who can sign in through your identity provider can get a Dot account automatically. While "Default new users to admin role" is on, each of those people becomes an admin. Turn it off first if you want new sign-ins to be regular users.
{% endhint %}

Make sure your IdP releases the `email` and (ideally) `name` claims to Dot. Most IdPs include both in the default profile scope; if yours doesn't, add `email profile` to the requested scopes on the application you registered in step 2.

## Troubleshooting

* **"redirect\_uri mismatch"** — the URL in your IdP's allowed callback list doesn't exactly match the Redirect URI shown in Dot. Re-copy it; watch for a missing trailing path or a leading `https://` typo.
* **Login button doesn't appear** — make sure all three required fields (Client ID, Client Secret, Metadata URL) are saved. The button is gated on a successful save, not just on form input.
* **"invalid\_client"** — the Client Secret is wrong or has been rotated in the IdP. Generate a new secret and re-save it in Dot.
* **Metadata URL fetch fails** — verify the URL works in a browser and returns a JSON document with `authorization_endpoint`, `token_endpoint`, and `jwks_uri` fields. If your IdP is on a private network, Dot's servers won't be able to reach it; expose the discovery endpoint publicly or use a named provider.

## Removing the integration

To stop accepting Generic OIDC sign-ins, click **Remove** on the tile. The button disappears from the login page immediately. Existing users provisioned via OIDC are kept and can still be reached if you re-enable any SSO provider — they aren't deleted with the integration.


# Security & Privacy

Ensuring high standards to protect customer data is critical to operate successfully given rising IT security threads and increased privacy concerns.

We are continuously auditing our technical and organizational measures by monitoring our infrastructure and processes with [Secureframe](https://secureframe.com/).

We've successfully completed the AICPA Service Organization Control (SOC) 2 Type I and Type II audit. The audit confirms that Snowboard Software GmbH’s information security practices, policies, procedures, and operations meet the SOC 2 standards for security. The audit was conducted by [Prescient Assurance](https://www.prescientassurance.com/).

For all questions and documentation requests, you can contact our customer support team at <hi@getdot.ai>.

<figure><img src="/files/Y115iHyah4R3SWmOAf6R" alt="" width="203"><figcaption></figcaption></figure>

## Organizational Security

**Information Security Program**

We have an Information Security Program in place that is communicated throughout the organization. Our Information Security Program follows the criteria set forth by the SOC 2 Framework. SOC 2 is a widely known information security auditing procedure created by the American Institute of Certified Public Accountants.

**Third-Party Audits**

Our organization undergoes independent third-party assessments to test our security and compliance controls.

**Third-Party Penetration Testing**

We perform an independent third-party penetration at least annually to ensure that the security posture of our services is uncompromised.

**Roles and Responsibilities**

Roles and responsibilities related to our Information Security Program and the protection of our customer’s data are well defined and documented. Our team members are required to review and accept all of the security policies.

**Security Awareness Training**

Our team members are required to go through employee security awareness training covering industry standard practices and information security topics such as phishing and password management.

**Confidentiality**

All team members are required to sign and adhere to an industry standard confidentiality agreement prior to their first day of work.

**Background Checks**

We perform background checks on all new team members in accordance with local laws.

## Cloud Security

**Cloud Infrastructure Security**

All of our services are hosted with Amazon Web Services (AWS). They employ a robust security program with multiple certifications. For more information on our provider’s security processes, please visit AWS Security.

**Data Hosting Security**

All of our data is hosted on Amazon Web Services (AWS) databases. These databases are located in the United States or Europe. Please reference the above vendor specific documentation linked above for more information.

**Encryption at Rest**

All databases are encrypted at rest.

**Encryption in Transit**

Our applications encrypt in transit with TLS/SSL only.

**Vulnerability Scanning**

We perform vulnerability scanning and actively monitor for threats.

**Logging and Monitoring**

We actively monitor and log various cloud services.

**Business Continuity and Disaster Recovery**

We use our data hosting provider’s backup services to reduce any risk of data loss in the event of a hardware failure. We utilize monitoring services to alert the team in the event of any failures affecting users.

**Incident Response**

We have a process for handling information security events which includes escalation procedures, rapid mitigation and communication.

Access Security

**Permissions and Authentication**

Access to cloud infrastructure and other sensitive tools are limited to authorized employees who require it for their role.

Where available we have Single Sign-on (SSO), 2-factor authentication (2FA) and strong password policies to ensure access to cloud services are protected.

**Least Privilege Access Control**

We follow the principle of least privilege with respect to identity and access management.

**Quarterly Access Reviews**

We perform quarterly access reviews of all team members with access to sensitive systems.

**Password Requirements**

All team members are required to adhere to a minimum set of password requirements and complexity for access.

**Password Managers**

All company issued laptops utilize a password manager for team members to manage passwords and maintain password complexity.

## AI Security

**Secure AI Model Usage**

[Dot](https://getdot.ai/) is powered by OpenAI’s and Anthropic’s leading frontier models. These models provide state-of-the-art AI capabilities while maintaining strong security and privacy measures.

**No Training on Customer Data**

Data sent to these models will never be used for training. For more details, refer to:

* [OpenAI’s API Data Usage Policies](https://openai.com/policies/api-data-usage-policies)
* [Anthropic’s Privacy Policy](https://privacy.anthropic.com/en/articles/7996885-how-do-you-use-personal-data-in-model-training?utm_source=chatgpt.com#h_1a7d240480)

As an organization, we are committed to safeguarding customer data and will never use it to train our AI models.

**AI Data Protection Measures**

* Encryption: All data exchanged with AI models is encrypted in transit using TLS/SSL to prevent unauthorized interception.
* Access Control: Only authorized systems and users can interact with AI processing services, ensuring strict data governance.
* Logging & Monitoring: AI interactions are logged and monitored for anomalies to detect and prevent misuse.

**Compliance & Risk Management**

* Third-Party Security Audits: AI security practices are included in our periodic security assessments and SOC 2 audits.
* Regulatory Alignment: AI data handling complies with GDPR, CCPA, and other applicable data protection regulations.
* Incident Response: If an AI-related security incident occurs, we have an established protocol for rapid investigation and resolution.

## Vendor and Risk Management

**Annual Risk Assessments**

We undergo at least annual risk assessments to identify any potential threats, including considerations for fraud.

**Vendor Risk Management**

Vendor risk is determined, and the appropriate vendor reviews are performed prior to authorizing a new vendor.

## Responsible Disclosure

At Dot, we consider the security of our systems a top priority. But no matter how much effort we put into system security, there can still be vulnerabilities present.

If you discover a vulnerability, we would like to know about it so we can take steps to address it as quickly as possible. We would like to ask you to help us better protect our clients and our systems.

Please do the following:

* E-mail your findings to <security@sled.so>,
* Do not take advantage of the vulnerability or problem you have discovered, for example by downloading more data than necessary to demonstrate the vulnerability or deleting or modifying other people's data,
* Use test.getdot.ai for security testing and not one of our production services,
* Do not reveal the problem to others until it has been resolved,
* Do not use attacks on physical security, social engineering, distributed denial of service, spam, automated probing that strains our servers, or applications of third parties, and
* Do provide sufficient information to reproduce the problem, so we will be able to resolve it as quickly as possible. Usually, the IP address or the URL of the affected system and a description of the vulnerability will be sufficient, but complex vulnerabilities may require further explanation.

What we promise:

* We will respond to your report within 3 business days with our evaluation of the report and an expected resolution date,
* If you have followed the instructions above, we will not take any legal action against you in regard to the report,
* We will handle your report with strict confidentiality, and not pass on your personal details to third parties without your permission,
* We will keep you informed of the progress towards resolving the problem,
* In the public information concerning the problem reported, we will give your name as the discoverer of the problem (unless you desire otherwise), and
* As a token of our gratitude for your assistance, we offer a reward for every report of a security problem that was not yet known to us. The amount of the reward will be determined based on the severity of the leak and the quality of the report. The minimum reward will be a 50 USD gift certificate.

We strive to resolve all problems as quickly as possible, and we would like to play an active role in the ultimate publication on the problem after it is resolved.<br>


# Support

Whenever, Wherever

You can use the following ways to contact support:

* Reach out by email to <hi@getdot.ai>
* Use the Support Chat integration in the app ❓
* Shared Slack Channel
* Shared Team Chat


# Core Concepts


# AI Agents

## What Exactly Is an AI Agent?

An **AI agent** is essentially an autonomous software system that uses artificial intelligence to [perceive its environment, make decisions, and take actions](https://en.wikipedia.org/wiki/Intelligent_agent) in pursuit of a goal. In simpler terms, an AI agent can be thought of as a program that acts on your behalf, **completing tasks without needing constant human direction**. These agents are built to exhibit capabilities like reasoning about how to achieve objectives, planning out sequences of steps, **adapting to new information, and learning from experience**. Unlike a standard software application that only follows predefined rules, an AI agent can **autonomously figure out what to do next** based on the situation – it has a degree of self-directed intelligence.

AI agents often incorporate **large language models (LLMs)** or other AI models as a core component of their “brains” to understand instructions and generate responses. But they go beyond a single model’s capabilities. A typical AI agent will maintain an internal **memory** of context or past interactions, and it can [**interface with tools or external systems**](https://www.anthropic.com/engineering/building-effective-agents) (like databases, APIs, or even physical devices) to get information or execute actions in the real world. In other words, the agent perceives inputs (from users or the environment), processes them to decide on an action, and then **acts through effectors or tools** to change something or produce an outcome. This loop of **perception, reasoning, and action** repeats as the agent works step-by-step towards its goal.

<figure><img src="/files/gmwnQ7yFRLCROx7Yjqtq" alt=""><figcaption></figcaption></figure>

*This diagram shows how an AI agent ingests user input and environment data, then reasons and plans using internal memory, calls external tools or services, and produces an action or response.*

An AI agent can be something as simple as a **basic chatbot** that responds to questions, or as complex as an **autonomous vehicle’s driving system** making real-time decisions on the road. The defining trait is that the agent has **some autonomy**: it doesn’t need to be told every little step, and it can **operate within set boundaries to achieve a goal**. Modern AI agents are capable of multi-step reasoning – for example, breaking down a complex request into sub-tasks and deciding which tools to use or what actions to take in sequence. They’re not limited to just chatting; they can execute code, look up information, control Internet-of-Things devices, or whatever their toolset and permissions allow. In short, an AI agent is **software that can “think and act” in a limited domain on its own**, using AI to figure out the *how* and *when* of completing the tasks we delegate to it.

## Why Do Organizations Need AI Agents?

Organizations are turning to AI agents as a way to **work smarter, faster, and at greater scale**. Unlike traditional software that might automate a fixed process, AI agents can handle **complex, dynamic tasks that were previously hard to automate** with prewritten rules. This means businesses can offload more challenging or time-consuming tasks to agents, allowing human teams to focus on higher-level strategy and creativity.

One major reason for adopting AI agents is **increased productivity and efficiency**. Autonomous agents can take over the “busy work” and the constant decisions required in complex workflows, operating tirelessly and quickly. For example, an AI agent could monitor and adjust thousands of cloud servers or perform continuous market research updates without any human intervention. This 24/7 capability **boosts overall throughput**, as the agent doesn’t need breaks and can respond to events or requests at any hour. In enterprise settings, AI agents have been shown to **save teams time by handling multi-step processes automatically**, drastically reducing the manual effort needed for things like data analysis, customer support inquiries, or software testing. A well-implemented agent can complete tasks in minutes that might take employees hours, all while maintaining a high level of accuracy.

Accuracy and consistency are another benefit. Because AI agents can be designed to **self-check their work and correct errors**, they often produce highly reliable results. For instance, an agent drafting financial reports or performing quality control checks can catch anomalies or mistakes that humans might miss when fatigued. These agents can continuously refine their outputs via feedback loops – essentially learning from mistakes – which helps organizations reduce errors in processes and make outcomes more predictable.

AI agents also help organizations **be more agile and responsive**. In a fast-changing environment (be it market conditions, supply chain disruptions, or cybersecurity threats), agents can react in real time or even **proactively anticipate problems**. A proactive agent might detect a subtle shift in data – say, a pattern that indicates a potential fraud or a looming equipment failure – and take action or alert staff before the issue grows. In areas like IT operations or security, such agents act as ever-vigilant sentinels, scaling up an organization’s ability to respond instantly to incidents.

Moreover, AI agents can **fill skill gaps and amplify expertise**. Not every company has an expert on every subject in-house, but an AI agent with the right training or access to knowledge can emulate some of that expertise. For example, if a company lacks a data science team, an AI agent could analyze complex datasets and surface insights that drive decision-making. In a world where certain skills are in high demand and short supply, these agents become force multipliers – they allow a small team to punch above its weight by tackling specialized tasks (from coding to customer engagement) that otherwise might require additional hiring or contracting.

Other organizational needs met by AI agents include **cost savings** and **scalability**. Automating tasks through agents can be dramatically cheaper in the long run than paying for extensive manual labor, especially for routine but cognitively heavy tasks. While there’s an upfront investment in setting up an agent, the marginal cost of its work is low – it can operate as an always-on employee that doesn’t draw a salary. And as the business grows, **agents can easily scale** to handle more volume by simply running more computations, something humans can’t do without a proportional increase in headcount. This scalability means organizations can handle surges in work (like seasonal customer queries or end-of-quarter data processing) smoothly with the help of AI agents.

Finally, AI agents often **break down silos in organizations** by linking processes and data across departments. Because one agent (or a team of agents) can be given access to multiple systems, they can coordinate tasks that span traditionally separate functions. Imagine an agent that automates an order fulfillment process: it could pull data from sales records, trigger actions in the warehouse system, coordinate with a finance system for invoicing, and update a customer service platform – all in one continuous workflow. This ability to traverse systems means processes become more integrated, and the flow of information improves across the organization.

In summary, organizations adopt AI agents to **enhance efficiency, accuracy, availability, and scalability** in their operations. By entrusting autonomous AI with complex tasks, businesses can reduce manual workloads, respond faster to opportunities or issues, cut costs, and leverage advanced insights – all of which are significant competitive advantages in today’s data-driven environment. AI agents essentially enable a level of **automation plus intelligence** that goes beyond what earlier automation or business software could do, addressing the growing need for smarter workflows in enterprises.

## Different Types of AI Agents

Not all AI agents are alike – in fact, there’s a broad spectrum of agent types, each suited to different scenarios. At one end, we have **simple reactive agents**, sometimes called reflex agents. These follow straightforward *if-this-then-that* rules: they respond to the current input or environment state with a predetermined action, without much internal thinking. A classic example would be a basic thermostat (in an AI context) or a rule-based chatbot that always replies with a canned response to certain keywords. Reactive agents tend to lack memory of past events; they operate only on the here-and-now. Because of this simplicity, reactive agents are **very reliable within their narrow scope** – they do exactly what they’re programmed to do – but they can’t handle situations outside those preset rules.

Moving up in sophistication, we encounter **proactive agents**. These agents don’t just wait for a prompt; they can **anticipate needs or changes** and act on their own. Proactive AI agents often utilize predictive algorithms and pattern recognition to make informed guesses about what might happen next. For instance, a proactive agent monitoring a supply chain might notice trends that suggest a shipment delay and proactively reroute orders or inform managers before anyone asks. In essence, proactive agents have a degree of *initiative* – they strive to achieve goals set for them by looking ahead, not just reacting to the immediate state of the world.

There are also **hybrid agents** that combine reactive and proactive strategies. Think of these as agents with a “best of both worlds” approach: they can handle routine situations with efficient, rule-based responses, but when the context becomes complex or unanticipated, they switch into a more deliberative, predictive mode. Many practical AI systems in business are hybrid. For example, a customer service AI might immediately handle common requests with scripted answers (reactive), but if it detects a frustrated customer or an unusual query, it might engage a more sophisticated problem-solving mode or escalate to a human (a kind of proactive decision).

Another category to note is **utility-based agents**. These agents are driven by a goal of maximizing some “utility” value – essentially, they evaluate different possible actions and choose the one that scores highest according to a utility function (a measure of success or happiness). Utility-based agents try to **optimize outcomes rather than just achieve a simple goal**. A self-driving car’s navigation system can be seen as utility-based: it evaluates routes and picks the one that minimizes travel time and risk, effectively maximizing the utility of a quick, safe arrival. These agents handle trade-offs and complex criteria well, because they are always asking “what’s the best action I can take now for the best overall result?”

We also have **learning agents**, which, as the name suggests, improve themselves through experience. A learning agent might start with some basic knowledge or behavior and then adjust its strategies by observing how well they work. It has components to gather feedback (from the environment or from user corrections) and update its internal models. Over time, a learning agent becomes more competent in its domain. For example, consider an AI personal assistant that learns a user’s preferences over time – initially it might make generic suggestions, but as it learns from the user’s reactions and corrections, it starts tailoring its actions (like scheduling meetings or prioritizing emails) in a way that better suits the user.

Finally, in enterprise contexts we’re seeing **collaborative agents** and multi-agent systems. Rather than a single agent trying to do everything, you have a *team of agents* that specialize and work together. Collaborative AI agents coordinate their actions, communicate with each other, and even negotiate responsibilities to solve complex tasks across different domains or departments. You might have one agent that specializes in gathering data, another that analyzes it, and a third that takes action on the analysis – collectively working on a larger process. They can also include humans in the loop as part of the network of collaboration. This mirrors how human teams operate, and it’s powerful for tackling big problems. Collaborative agents break down silos by **coordinating across different functions**, sometimes described as a digital workforce working in concert.

It’s worth noting that these categories aren’t mutually exclusive – a single AI agent might be both learning and utility-based, for example. But this taxonomy gives a sense of the range from the most constrained agents to the most sophisticated. Understanding the type of agent helps organizations set the right expectations: a reactive agent will be reliable for repetitive tasks but won’t handle surprises, whereas a learning, utility-driven agent might handle surprises but needs monitoring and training. Selecting the right type (or mix) for the task at hand is a key part of adopting AI agents effectively.

## What Is the Difference Between LLMs, AI Agents, and AI Workflows?

As AI has exploded in popularity, a lot of terms like *LLM*, *agent*, and *workflow* get used interchangeably or without clear meaning. In reality, they refer to related but distinct concepts in the AI ecosystem. Let’s break them down:

* **LLM (Large Language Model):** This typically refers to a machine learning model (often based on deep neural networks) that’s trained on a vast amount of text data to understand and generate human-like language. An LLM on its own is **not an agent** – it’s more like a very advanced predictor of text. It takes an input (a prompt) and produces an output (a continuation or answer) based on patterns it learned from training data. Importantly, an LLM by itself doesn’t have goals or take actions *beyond* producing a coherent response to each prompt. It doesn’t remember anything beyond what you give it (aside from its internal learned parameters), and it won’t spontaneously decide to do something without being asked. Think of an LLM as the “brains” in terms of linguistic ability – it can reason and draw on knowledge to answer a question, but it’s passive in that it only acts when prompted, and its action is just generating text.
* **AI Agent:** An AI agent, as we’ve discussed, is an autonomous system that can *use* models like LLMs as components, but it also has the machinery to make decisions and take actions toward a goal. The key difference is **autonomy and agency**. An AI agent typically will have an LLM or similar model under the hood, but it wraps that model with additional logic: the ability to plan a series of steps, the ability to call external tools or APIs, and the ability to adjust its plan based on what happens at each step. So, while an LLM might answer a single question in isolation, an AI agent could be given an objective like “Schedule meetings for next week with all team leads” and then figure out how to achieve it – perhaps by querying calendars, sending emails, and using the LLM to draft polite meeting requests. The agent orchestrates these actions **without needing the user to prompt each step**. In essence, the agent leverages the intelligence of models (like LLMs) but adds a layer of **decision-making and action-taking** that gives it independence.
* **AI Workflow:** An AI workflow (sometimes called an AI automation or simply an automated workflow with AI elements) lies somewhere between traditional automation and an agent. An AI workflow is usually a **predefined sequence of tasks or steps**, some of which might involve AI. For example, imagine a customer support process that is set up as a workflow: first, classify the incoming ticket (maybe using an AI model to read the text), then send an automated reply or route it to the appropriate team, then create a summary report at the end of the week. Each of those steps can be laid out explicitly by a programmer or using a workflow automation tool. The AI comes into play perhaps in classifying the text or generating a summary, but the **flow of control is fixed** – the system isn’t deciding on its own to add new steps or change the order; it’s following a script that just happens to call AI models for certain tasks.

<figure><img src="/files/N6FuMkFNGlVTtygTdmyZ" alt=""><figcaption></figcaption></figure>

*Diagram: Shows an AI workflow that chains LLM calls together in a program*

The confusion often arises because an AI agent might internally execute a dynamic “workflow” of its own creation, and conversely, an AI workflow can appear somewhat agent-like if it’s complex. The fundamental distinction is **who (or what) is driving the sequence of actions**. In an AI workflow or automation, a human designer has laid out the game plan (even if it branches conditionally), and the system just follows it. It won’t do anything that’s not in the plan. By contrast, an AI agent is given an end goal and it *decides* the sequence of actions on the fly – possibly even coming up with actions that weren’t explicitly anticipated by its creators (within the bounds of its capabilities).

Another way to look at it: **automation vs. agentic AI is about flexibility and decision autonomy**. A simple automation (or workflow) is great for repetitive, well-defined tasks and is highly predictable – it will do exactly the same thing every time unless we reprogram it. An AI agent is designed for open-ended tasks where we *want* the AI to figure out the procedure because it’s too complex or varied to script every scenario.

In summary, **LLMs are the intelligence engines (especially for language) but lack agency**, **AI workflows are streamlined automations with fixed logic that can include AI components**, and **AI agents are autonomous problem-solvers that dynamically plan and act, often using LLMs and other tools along the way**.

## What’s the History of AI Agents?

The concept of AI agents may seem cutting-edge, but it actually builds on decades of ideas and developments in artificial intelligence. Going back to the **1950s and 1960s**, AI pioneers like Alan Turing and others were already wrestling with what it would mean for a machine to act intelligently (Turing’s famous question, “Can machines think?”, led to the [Turing Test](https://plato.stanford.edu/entries/turing-test/) in 1950). Early programs in this era, however, were not agents by today’s definition – they were often chess-playing programs or theorem provers that demonstrated logic and reasoning in narrow domains. The notion of an “agent” – something that perceives and acts in an environment – was more formally introduced later.

By the **1970s and 1980s**, AI had moved into the era of **rule-based expert systems**. These systems, like MYCIN for medical diagnosis, could be seen as precursors to agents: they would take input (symptoms, for example) and produce actions (diagnostic recommendations). However, they lacked autonomy in the sense that they didn’t operate continuously or decide to act; they were essentially consultative programs. Researchers did start discussing *intelligent agents* around this time, at least conceptually. In 1988, a significant breakthrough came in **reinforcement learning** (a technique where an agent learns how to act through trial and error rewards). Reinforcement learning put the formal idea of an “agent” on the map in academia – here you had software agents in simulations learning to make sequences of decisions (like a game player improving its strategy).

The **1990s** saw the term **“intelligent agent” gain popularity** in research and even some early consumer software. This was when people started envisioning software that could act on behalf of a user in a networked environment. Academic projects and conferences on multi-agent systems blossomed. For example, researchers worked on agents for information retrieval that could roam the early internet to find information for you, or personal digital assistants that had a bit more smarts. One iconic (if not entirely successful) example from the late ’90s was [**Microsoft’s Clippy**](https://en.wikipedia.org/wiki/Office_Assistant), which tried to act as a little agent that noticed what you were doing

and offered help. Clippy was rule-based and not very well-loved, but it was a sign of the times – the dream was to have proactive helper agents. In the late 90s, **autonomous robot agents** were also being explored: robots that could navigate environments (like robotic vacuums, an early version of which, the iRobot Roomba, came out in 2002 as a simple cleaning agent for your home).

On the academic side, the 1990s also gave us frameworks for agent architectures – terms like **“belief-desire-intention” (BDI) agents** emerged, which tried to formalize what it means for an agent to have certain beliefs about the world, desires (goals), and intentions (plans it’s committed to).

Jump to the **2000s**, and we see machine learning (especially statistical learning) start to improve AI capabilities. **Agents began to incorporate machine learning**, making them less rigid. A big milestone was [**IBM’s Watson**](https://www.ibm.com/watson) winning *Jeopardy!* in 2011 (development started mid-2000s). Watson wasn’t exactly an autonomous agent living in the wild – it was a question-answering system – but it combined many AI techniques and could be seen as an agent that “reads” a clue (perceives) and “buzzes in with an answer” (acts) based on confidence. Around the same time, we saw the rise of **virtual assistants like Siri (2011)**, **Google Now (2012)**, and [**Cortana**](https://en.wikipedia.org/wiki/Cortana_\(Halo\)) **(2014)**. These were important because they introduced AI assistants to consumers, showing that software could listen to your voice, understand your request, and perform tasks (like setting reminders or searching information). These assistants were primitive agents: they were mostly reactive and followed scripted workflows for each command, but they had elements of natural language understanding and could interface with tools (phone apps, web search, etc.).

The **2010s** is when AI agents really started to get powerful, thanks in large part to advances in **deep learning** and especially the later half with **deep reinforcement learning** and **transformer-based language models**. In 2012, deep learning showed its prowess in perceiving the world (image recognition breakthroughs), which trickled into better computer vision for agents. By the late 2010s, [**DeepMind’s AlphaGo (2016)**](https://www.youtube.com/watch?v=WXuK6gekU1Y) demonstrated an agent that could learn to play Go at superhuman level, combining planning and learning. Though a game agent, this showcased how far autonomy had come (it could beat humans in a very complex task by training itself). On the language side, OpenAI’s [**GPT-3**](https://openai.com/index/gpt-3-apps/) **(2020)** was a watershed moment for LLMs – suddenly AI could generate text that was often indistinguishable from human writing. This directly paved the way for more conversational and capable agents.

Which brings us to **today (2020s)**: sometimes called the era of **“Agentic AI”**. Now all the pieces (perception, language, learning, planning) are coming together. Agents today can use off-the-shelf APIs for vision (to see), powerful pretrained models to understand and generate language (to reason and communicate), and have access to the knowledge on the internet or big data in real-time. The year 2023 in particular saw a surge of interest in truly autonomous agents when projects like [**AutoGPT**](https://github.com/Significant-Gravitas/AutoGPT) and others went viral – these are essentially LLM-powered agents that can recursively prompt themselves, spawn new sub-agents, and attempt to carry out open-ended goals.

We’re also seeing **multi-agent systems** being deployed in practical uses – for example, some advanced business solutions might have an “AI workforce” of multiple agents handling different tasks and passing information among themselves, simulating a team. In the physical world, self-driving cars and drones are essentially autonomous agents, continuously perceiving (through sensors) and acting (steering, etc.) to reach a destination safely.

So, from the **simple rule-based programs** of decades past to the **learning, conversational and tool-using agents of today**, the evolution has been one of increasing autonomy and sophistication. Each era brought a key capability: rules gave them structured knowledge, machine learning gave them adaptability, deep learning gave them perception and language, and now large models + tool use give them a form of open-ended cognitive flexibility. It’s been a journey of expanding what agents can do on their own.

## How Do AI Agents Fit Into Automation and Analytics?

AI agents can be seen as the next evolution in the automation landscape. In traditional **business process automation**, you might have a sequence of steps automated by scripts or robotic process automation (RPA) bots – but those are usually rigid and rule-based. AI agents step in to **add adaptability and decision-making to automation**. Rather than just following a fixed script, an agent can decide *when* to invoke certain automations or how to handle exceptions.

In the realm of **data analytics**, AI agents are providing a huge leap in how insights are generated and consumed. Traditionally, to get insights from data, you needed analysts to write queries, produce reports, and interpret results. Now imagine an AI agent that serves as a **virtual data analyst** – this is already happening. Such an agent can understand a business user’s question in natural language, translate it into the necessary database queries, fetch the data, perform analysis, and even generate visualizations or a written summary of findings. This is essentially what [Dot, the AI data analyst](https://getdot.ai) agent is designed to do: it connects to your data warehouse and business intelligence systems so that any team member can ask a question like “What were our top-selling products last quarter and why?” and get an answer with charts and explanations. Dot “digs through enterprise data” by using an LLM to interpret the question and figure out which data is relevant, then it may compose an SQL query to run on a Snowflake or BigQuery database, then take the results and present an insight – all in a conversational manner. In fact, as the Dot team describes, the agent even pulls in metadata (like data schema, documentation, query history) to make sure it’s interpreting things correctly, showing that it acts like a savvy human analyst who knows where to find the data and how to join it.

The value of such agents in analytics is enormous: **they empower non-technical users to explore data** on their own and get immediate answers, rather than waiting days for a data team to provide a report. This democratization of data analysis can make organizations much more responsive and data-driven in everyday decisions.

AI agents in analytics also routinely handle tasks like **monitoring data quality and detecting anomalies**. Data observability companies are beginning to embed AI that acts somewhat like an agent: it watches data pipelines for unusual patterns and then alerts teams or even initiates corrective steps. These agents ensure that the data feeding business decisions is reliable, effectively automating part of the data engineering oversight.

Moreover, agents tie into the analytics ecosystem by integrating with existing tools: for instance, an agent might use a visualization tool’s API to automatically generate a dashboard when asked, or it might interface with collaboration tools like Slack/Teams to deliver analytics results to users where they already communicate. Notably, [Dot can live in Slack or Teams chat](https://getdot.ai) – meaning the agent is embedded in the organization’s communication channels, ready to answer data questions or perform analyses on the fly.

To illustrate, [Dot is described as “an agent that works for you”](https://getdot.ai) – you can ask it in plain language about, say, sales figures or even instruct it to do something like “if our daily revenue falls below X, create a Jira ticket for the engineering team”. In that scenario, Dot moves beyond analysis to action, effectively automating a response to an analytic insight (it noticed something and created a task in another system). This blurs the line between analytics and operations automation, which is precisely how AI agents add value: they *span multiple domains of work* seamlessly.

In summary, AI agents complement existing automation by making it smarter and more adaptable, and they revolutionize analytics by making data interaction conversational and proactive. They **fit into the overall ecosystem as a connective tissue** – linking data to action. Businesses get the benefit of automation (speed and consistency) combined with a form of intelligence (contextual understanding and adaptability) that was missing before.

## What Are Typical Use Cases for AI Agents?

AI agents are being applied across a wide range of industries and functions – essentially anywhere there are complex, repetitive, or information-intensive tasks, an autonomous agent can potentially help. Here are some prominent use cases and examples:

* **Customer Service and Support:** AI agents in customer service go beyond static chatbots. They can handle entire customer interactions, from understanding the query to taking action. For instance, an AI agent might handle a refund request: it can converse with the customer to gather details, authenticate the customer’s identity, process the refund in the backend system, and confirm the resolution – all without a human agent’s involvement. These agents use natural language understanding to parse customer emails or chats and often hook into CRM systems, order databases, etc., to get the job done. They **learn from each interaction**, so over time they get better at resolving issues.
* **Software Development and IT Operations:** AI agents are starting to act as co-developers or DevOps assistants. For example, developers can task an AI agent with writing simple functions or even entire microservices. In IT operations, agents can monitor infrastructure – there are agents that watch logs and metrics, detect anomalies, and then act, such as by restarting services or applying patches.
* **Finance and Banking:** In finance, AI agents are used for things like **fraud detection**, portfolio management, and algorithmic trading. A fraud-detection agent can scrutinize transactions in real-time and if it sees a pattern that looks suspicious, it can autonomously flag the account, halt transactions, or even initiate further verification steps. Banks also deploy AI agents as virtual financial advisors or customer assistants.
* **Marketing and Sales:** AI agents in marketing can personalize customer outreach at scale. Imagine an agent that manages an email campaign – it could automatically segment customers based on their behavior, craft tailored content for each segment, send out the emails at optimized times, and then analyze the response. More sophisticated agents might handle inbound sales inquiries, qualify leads, or even schedule demos.
* **Human Resources:** Companies are experimenting with AI agents to streamline HR tasks. One use case is in **recruiting** – an AI agent can conduct an initial interview with candidates through a chat or email, then evaluate the answers to short-list candidates for the human recruiters. In employee onboarding, an AI agent might guide a new hire through all the setup steps: completing forms, learning about company policies, and ensuring they have access to all needed systems.
* **Healthcare:** We have virtual health assistants that can interact with patients to collect symptoms and give basic guidance. Some agents assist doctors by auto-populating medical records or monitoring patient follow-up.
* **Supply Chain and Manufacturing:** AI agents here can monitor supply chain events end-to-end and take action to optimize flow. For example, if an agent sees that a shipment is delayed due to weather, it could proactively reroute other shipments or adjust factory production schedules.

The World Economic Forum has highlighted applications from **software development to education to finance to customer service to healthcare**, showing that AI agents are a general paradigm that can specialize in many tasks. In summary, typical use cases for AI agents include **any repetitive decision process, any situation requiring continuous monitoring, and any service that can be made more efficient by personalization or rapid reaction**.

## What Should You Look Out For When Using or Buying an AI Agent?

Deploying AI agents in an organization isn’t just a plug-and-play affair – there are important considerations to keep in mind to ensure you get the value you want without unintended consequences.

* **Autonomy and Control:** First, determine how much autonomy the agent will have and how you will control it. Not all agents are fully autonomous; some operate in a “human-in-the-loop” mode where they make recommendations but a person approves the final action. It’s wise initially to set the agent’s permissions narrowly and always have a **manual override or an off-switch**.
* **Transparency and Explainability:** One big challenge with AI agents (especially those using complex models like LLMs) is that their decision-making can be a black box. When the agent takes an action, you’ll want to know **why** it did that. At minimum, ensure the agent’s actions are **logged and auditable**. This is crucial not just for trust, but for debugging when things go wrong.
* **Data Privacy and Security:** AI agents often need access to a lot of data and systems to be effective – they might connect to your databases, customer data, internal APIs, etc. That raises questions: Is the agent handling data securely? Is sensitive data protected? When evaluating a vendor, check their security certifications and how they handle your data. Ideally, an AI agent for enterprise use should allow deployment in a secure cloud or on-premises if needed, with encryption of data in transit and at rest. Also consider **access control**.
* **Accuracy and Reliability:** AI agents can and will make mistakes. They might misunderstand a request, use the wrong tool, or get an analysis wrong. When adopting an agent, it’s important to **benchmark and test its performance** on tasks that matter to you. You should plan for a **pilot phase** where the agent’s outputs are closely monitored by staff.
* **Ethical and Compliance Considerations:** AI agents acting on behalf of your organization must adhere to all the laws, regulations, and ethical norms that a human employee would. This includes things like **avoiding bias** in decisions, respecting customer privacy preferences, and maintaining appropriate tone and fairness in communications. Organizations should **establish clear ethical guidelines and guardrails for AI agents**, including human oversight for sensitive tasks.
* **Misalignment and Unintended Actions:** An AI agent could, if not properly constrained, do things that are technically within its goal but not actually desirable (the classic “alignment problem”). To avoid this, define the agent’s objectives carefully and include multiple success criteria.
* **Vendor Promises vs Reality:** If you are looking to buy an AI agent solution, be a bit skeptical of marketing claims. As with any hype, some products labeled “AI agents” might just be glorified scripts or chatbots. Ask the vendor to clarify what exactly their agent can do autonomously.
* **Integration and Maintenance:** Consider how the AI agent will integrate with your existing systems. Does it have connectors or APIs for the software you use? Also, think about maintenance – AI models need updating. Maintenance also includes monitoring the agent’s ongoing performance.

In essence, using or buying an AI agent responsibly means **treating it not as a magic box, but as a powerful tool that needs governance**. Establish checks and balances: for example, have a process for employees to flag if the agent made a bad call, and a way to quickly correct or update it.

## Where Are AI Agents Headed?

AI agents today are impressive, but the trajectory of their development suggests they will become even more capable, ubiquitous, and integrated into our lives and businesses. Looking to the future, several key trends and possibilities stand out:

**1. Greater Collaboration (Multi-Agent Ecosystems):** We can expect to see more **multi-agent systems**, where many specialized agents work together (and with humans) as teams. Standards and protocols for **agent-to-agent communication** will likely advance so that heterogeneous agents can talk to each other securely and efficiently.

**2. Enhanced Abilities via Multimodal AI:** Right now, many AI agents are heavily text-based. In the near future, agents will be much more **multimodal** – meaning they can process and generate not just text, but images, audio, video, and other data forms.

**3. More Natural Interaction & Personalization:** As agents get better language models and interface design, interacting with them will become more intuitive. We’re heading towards agents that **remember your preferences and context over long periods**, leading to a truly personalized experience.

**4. Integration with Physical World (Robotics):** Many current AI agents live in the digital realm, but the line between digital and physical is blurring. Drones, robots, self-driving vehicles – these are physical embodiments of AI agents.

**5. Better Reasoning and Reduced Hallucinations:** Research is actively addressing LLM-driven agent limitations with techniques like chain-of-thought prompting and hybrid systems that combine neural networks with symbolic reasoning.

**6. Stronger Ethical and Safety Guardrails:** As agents become more powerful and autonomous, there will be an even greater emphasis on ensuring they act safely and in alignment with human values. In practice, future AI agents might have built-in “ethical modules” – subroutines that constantly evaluate the agent’s planned actions against ethical rules or policies.

**7. Wider Adoption and New Business Models:** Just as every company today has a website or an app, we might reach a point where **every company has AI agents as part of its workforce**. This means the tools to create and manage agents will become more user-friendly and widespread.

**8. Closer Human-AI Collaboration:** In the foreseeable future, rather than AI agents completely replacing humans, the trend is toward **collaboration** – what some call “augmented intelligence.” Agents will handle the heavy lifting of data processing, routine action, and first-level decisions, while humans will focus on oversight, strategic decisions, and tasks that involve high levels of empathy, creativity, or complex judgment.

**9. New Frontiers – Creativity and Emotional Intelligence:** Agents might develop a form of artificial emotional intelligence – that is, becoming better at reading human emotions and responding appropriately. The line between a “bot” and something that feels like a genuine personality could blur (raising its own ethical questions, of course, about transparency that it’s an AI).

In summary, the future of AI agents is one where they are **more everywhere, more capable, and more integrated into the fabric of work and daily life**. They will handle more of the drudgery and even some sophisticated tasks, while humans will supervise and focus on what humans do best. Achieving this future means tackling the challenges of today (alignment, safety, reliability) head-on, but the momentum in research and industry suggests we will. As these agents proliferate, it will be crucial to keep the human-centric perspective: using them to **amplify human potential and solve problems**, while steering clear of pitfalls. If done responsibly, AI agents are poised to become invaluable collaborators in virtually every field, helping us achieve things faster and perhaps tackle problems that were once deemed too complex to manage.


# AI Explainability

## What Exactly Is AI Explainability?

**AI explainability** (often called *explainable AI* or XAI) refers to techniques and methods that make the workings of artificial intelligence systems clear and understandable to humans. In essence, explainability is about answering *“Why did the AI make this decision?”* in human terms. Modern AI models – especially complex machine learning models like deep neural networks – are often **“black boxes”** whose internal logic is opaque, meaning even their designers can’t easily explain how inputs are being turned into outputs. Explainable AI aims to **counter this black-box problem** by providing insight into the AI’s reasoning process. This might involve highlighting which features of the input data influenced a prediction, providing simplified rule-based explanations, or offering visualizations that trace the model’s decision path.

The goal of AI explainability is not only to satisfy curiosity – it’s fundamentally about **human trust and oversight**. By making AI’s reasoning transparent, explainability lets users and stakeholders comprehend and **trust the results** that an AI system produces. An explainable model can articulate what it is doing, what information it is relying on, and why, giving humans the confidence that the AI’s decisions are sound. In practice, explainability techniques can improve the **user experience** of AI-powered products and services by assuring people that the AI is “making good decisions”. Even in cases where no law explicitly requires it, providing explanations for AI decisions can help users feel more comfortable and in control when interacting with an AI system. In short, AI explainability is about bridging the gap between complex algorithmic logic and human understanding, turning inscrutable computations into relatable insights.

## Why Do Organizations Need AI Explainability?

As AI systems play a growing role in business and society, organizations are recognizing that **explainability is essential, not optional**. There are several compelling reasons why companies and institutions need to make their AI models explainable:

**1. Building Trust and Adoption:** If people don’t understand or trust an AI’s decisions, they won’t use it – no matter how accurate it might be. Explainability is the foundation for trust in AI systems. Customers, employees, and other stakeholders need confidence that AI-driven recommendations or decisions are fair, sensible, and reliable. For example, a sales team is far more likely to follow an AI-generated recommendation if they’re given a clear reason *why* the AI suggests it, rather than if the suggestion comes out of a black box. Indeed, knowing the rationale behind an AI’s recommendation **increases users’ confidence** in acting on it. In high-stakes domains like finance or healthcare, trust is even more critical – a doctor or loan officer must be able to justify the AI’s decision to a patient or client. By providing human-readable explanations, organizations ensure that AI tools are actually adopted and used to their full potential, rather than dismissed due to a “magical” output that no one can vet.

**2. Improving Model Performance and Accountability:** Explainability isn’t just for end-users; it’s also a powerful tool for data scientists, engineers, and **MLOps teams** who monitor and refine AI models. Being able to peek under the hood of a model can significantly **boost productivity** in model development and maintenance. For instance, if an AI model makes an unexpected prediction, an explanation can reveal whether the model picked up on spurious correlations or data errors. Techniques that enable explainability can quickly highlight **errors or areas for improvement** in a model’s behavior, helping teams diagnose bugs or biases in the system. Understanding which input features most influenced a model’s output lets engineers verify that the model is learning the right patterns rather than latching onto noise. In this way, explainability acts as a debugging and validation aid: it provides an **audit trail** of how the AI reached its conclusion, so that developers can trace issues and ensure the model is behaving as intended. This accountability is vital for continuously improving AI systems and preventing subtle problems from going unnoticed.

**3. Surface New Insights and Business Value:** Sometimes, understanding *why* a model made a prediction can be as valuable as the prediction itself. Explanations can uncover actionable insights that would otherwise remain hidden in a black box. For example, suppose an AI model predicts a certain group of customers is likely to **churn** (cancel a service). That prediction alone is useful, but an explanation of **why** those customers might churn (perhaps due to specific product issues or price sensitivity) is even more powerful – it points to concrete interventions the business can take to reduce churn. In one case, an auto insurance company applied explainability tools (specifically using [SHAP values](https://github.com/slundberg/shap), a popular explainability technique) to its risk model. The explanations revealed that certain interactions between driver characteristics and vehicle features were driving up risk predictions – insights that weren’t obvious from the raw model output. By adjusting the model and underwriting policies based on those explanations, the insurer significantly improved its performance. This example illustrates how explainability can lead companies to **better decisions and strategies**: by illuminating the drivers behind model outputs, organizations can discover new levers for business value that a mere prediction wouldn’t show.

**4. Ensuring Alignment with Business Goals:** Organizations deploy AI with specific objectives in mind – but complex models don’t always behave as expected. Explainability helps business teams confirm that an AI system’s reasoning aligns with business logic and values. When technical teams can explain how an AI makes decisions, business stakeholders can verify that the model is optimizing for the right outcomes (for example, maximizing long-term customer value rather than short-term tricks) and that nothing was “lost in translation” between the business problem and the model’s mathematical goals. If an explanation reveals that a model is focusing on the wrong factors (say, a retail recommendation AI giving undue weight to irrelevant product attributes), the company can course-correct before the model causes real harm. In this sense, explainability acts as a **safety net** to ensure AI solutions truly serve the business purpose they were designed for, and it fosters better communication between data science teams and business units.

**5. Mitigating Risk and Meeting Regulatory Requirements:** Perhaps one of the most urgent drivers for AI explainability is risk management and compliance. As AI systems make decisions that affect people’s lives – deciding who gets a loan, a job interview, or an insurance policy – there is a growing demand for **accountability**. Regulators around the world have begun to insist on a “right to explanation” and algorithmic transparency in certain domains. In the financial industry, for instance, lenders in many jurisdictions must provide reasons to applicants who are denied credit. It’s no longer acceptable for a bank to say an algorithm rejected an application “just because” – they may need to point to specific factors like credit history or income that influenced the decision. In fact, some sectors *already require* explainability by law. A recent bulletin from the [California Department of Insurance](https://www.insurance.ca.gov/0400-news/0100-press-releases/2022/release117-2022.cfm), for example, mandates that insurers **explain any adverse actions** (like denying coverage or setting higher premiums) that are based on algorithmic models. And broader regulations are on the horizon: the [European Union’s proposed AI Act](https://artificialintelligenceact.eu/) includes explicit obligations for transparency and explainability in high-risk AI systems. Even when not explicitly mandated, providing explanations helps organizations avoid legal and ethical pitfalls. It allows internal risk and compliance teams to verify that an AI’s decisions **do not hide bias or discrimination**, and that they align with the company’s ethical standards and policies. In short, explainability is a key plank of **Responsible AI** – it helps prevent unintended harm by making the algorithm’s behavior visible and auditable. Companies that invest in explainable AI are essentially investing in protection against reputational damage, unfair outcomes, and compliance violations.

In summary, organizations need AI explainability to **build trust, drive adoption, improve their AI systems, unlock new value, and manage risks**. Studies have even found that companies getting the highest financial returns from AI are more likely to follow best practices for explainability. When people can understand and trust what an AI is doing, they are more likely to embrace it – and only then can its benefits be fully realized. As one report put it succinctly: *“People use what they understand and trust. This is especially true of AI.”* Organizations that prioritize explainability will not only satisfy regulators and avoid pitfalls, but also gain a competitive edge by using AI in a way that is transparent, accountable, and aligned with human values.

***

### Typical Explainability Engine Workflow

<figure><img src="/files/SYZbSpogzPAfleI7bpbA" alt=""><figcaption></figcaption></figure>

***

## A Brief History of Explainable AI

Although the term “explainable AI” has gained popularity in recent years, the pursuit of making AI systems understandable has deep roots in the history of artificial intelligence. In the early decades of AI (the 1970s–1990s), many AI systems were based on **symbolic reasoning** and expert rules, which were inherently more interpretable. For example, one of the earliest medical AI programs, **MYCIN** (developed in the 1970s to diagnose infections), relied on hand-coded if-then rules and could **explain which rules led to its diagnosis** in a given case. Similarly, expert systems like GUIDON and SOPHIE had built-in capabilities to articulate their problem-solving steps in a way a user or student could follow. These systems were limited in scope, but they demonstrated that AI could be made to **“think out loud”** by tracing through logic – an approach to explainability that was relatively straightforward when AI was essentially a collection of human-understandable rules.

The rise of machine learning, especially from the 1990s onward, shifted AI toward data-driven pattern recognition and complex statistical models. Models like neural networks and ensembles began to outperform rule-based systems, but they introduced a new challenge: their knowledge was encoded in numeric weights and connections, not human-readable rules. As early as the 1990s, researchers started asking whether it was possible to **extract explanations from trained neural networks**. In domains like healthcare, where new machine learning models were being developed to assist clinicians, there was a clear need to make these **opaque models more trusted and trustworthy** by providing dynamic explanations of their reasoning. The issue became more pressing in the 2010s as AI moved into high-stakes applications. Public concerns erupted over bias in algorithms used for things like criminal sentencing and credit scoring – for instance, news that a proprietary sentencing algorithm was biased against certain racial groups, or that a credit model was unfairly denying loans – which underscored the demand for **transparent AI**. These incidents prompted both academia and industry to develop tools that can detect and mitigate bias, and to push for algorithms whose decisions can be scrutinized and explained.

By the late 2010s, **Explainable AI (XAI)** had matured into a distinct research field. A notable milestone was the [DARPA XAI program](https://www.darpa.mil/program/explainable-artificial-intelligence) (launched in 2016–2017), which invested in new methods to produce “**glass box**” models – highly accurate machine learning models that are more **transparent** to human operators. The first international workshops and conferences on explainable AI took place around this time, reflecting a surge of interest in the topic. Researchers devised a variety of techniques: from **Layer-wise Relevance Propagation (LRP)**, which traces a neural network’s output back to the importance of each input feature, to local explanation methods focusing on individual predictions (often referred to as **“local interpretability”**). There was also renewed interest in “**glass box**” or interpretable models – like decision trees, generalized additive models, and sparse linear models – as ways to achieve high accuracy **with built-in explainability**.

Today, explainable AI remains a vibrant and fast-evolving area. Conferences such as [ACM FAccT](https://facctconference.org/) (Fairness, Accountability, and Transparency) dedicate attention to AI explainability in socio-technical systems, and an entire ecosystem of open-source libraries and commercial tools has emerged to help interpret complex models. We’ve come full circle in some respects: while early AI explained itself through explicit logic, modern AI often requires *post-hoc* explanation techniques to shed light on learned patterns. The difference now is scale and urgency – with AI systems affecting millions of people, the stakes for explainability are higher, and the techniques are far more sophisticated. From early expert systems to today’s deep learning models, the lesson remains clear: an AI that can explain its reasoning is crucial for human trust, effective collaboration, and ethical use of technology.

## Approaches to Explaining AI Models

How do we actually make an AI explain itself? There is no single answer – instead, there is a toolbox of **approaches to AI explainability**, each suited to different scenarios. Broadly, these approaches fall into two categories: **intrinsically interpretable models** and **post-hoc explanation techniques**.

* **Intrinsically interpretable models** are algorithms that are designed to be understandable from the start. These are sometimes called “**white-box**” models, in contrast to black-box models. For example, a simple decision tree is often interpretable because one can follow the tree’s splits (based on features) to see why a decision was made. Linear regression or logistic regression models are also considered interpretable – their decisions come from a weighted sum of features, so the weights can be examined to understand each feature’s influence. In fact, in some cases it’s possible to achieve high accuracy with such white-box models, negating the need for more complex algorithms. Domains like finance or healthcare often favor interpretable models for this reason. Researchers have even developed specialized inherently interpretable models (like generalized additive models with pairwise interactions, or “GA2M”) that try to match the accuracy of black-box models while remaining mostly transparent. The advantage of intrinsically interpretable models is that **explanation is built in** – you can usually point directly to the model’s structure (rules, weights, etc.) to explain its predictions, which simplifies governance and compliance.
* **Post-hoc explanation techniques** are methods applied *after* a complex model has been trained, in order to extract insights about its behavior. These are essential when using highly accurate but opaque models like random forests, gradient boosted machines, or deep neural networks. Post-hoc methods do not change the original model; instead, they analyze it from the outside. One common approach is to examine the model’s **feature importance** – essentially asking the model, *“which input features most affect your output?”* Many machine learning libraries can compute global importance scores (for example, by seeing how prediction error changes if a feature is shuffled or held out). However, global importance only provides a high-level view. To get more fine-grained explanations, especially for individual predictions, practitioners turn to techniques like [**LIME**](https://github.com/marcotcr/lime) and [**SHAP**](https://github.com/slundberg/shap):
  * **LIME (Local Interpretable Model-Agnostic Explanations)** creates explanations for a single prediction by perturbing the input and observing how the model’s output changes. In practice, LIME generates many slight variations of an input data point and uses the complex model to predict each variation; it then fits a simple, interpretable model (like a linear model) on those perturbations to approximate the complex model’s behavior *in the vicinity of that original data point*. The result is a small set of weights or rules that explain why the model made its prediction for that one instance. For example, if a neural network predicts that a certain customer will leave (churn), LIME might reveal a simple approximation like: *“if tenure < 1 year and support tickets > 3, then churn=Yes”* as a local explanation, indicating those factors drove the prediction. LIME is model-agnostic, meaning it can work with any type of classifier or regressor, and it’s been applied to explaining everything from text classifiers to image recognizers by focusing on parts of the input (like specific words or image segments).
  * **SHAP (Shapley Additive Explanations)** takes a game-theoretic approach to explanation. SHAP assigns each feature in a prediction an importance value – often called a **Shapley value** – that represents how much that feature contributed to the difference between the model’s actual prediction and some baseline prediction (such as the average over the dataset). The concept comes from cooperative game theory: imagine the model’s prediction is a “payout” and the features are players who each contribute to that payout. SHAP values are calculated in a way that **fairly distributes credit (or blame) among features** for the prediction. One intuitive way to think of it: SHAP considers all possible combinations of features and how adding a feature changes the model’s output, averaging these contributions in a principled manner. The end result is a set of feature attributions that sum up to the model’s prediction. For instance, a SHAP explanation for a house price prediction might say: *“Baseline price $200K + 50K (if `location = beachside`) + 30K (if `size = 2000 sqft`) – 10K (if `old roof`) = $270K predicted price.”* Such an explanation shows the *direction* and *magnitude* of each feature’s influence. SHAP is powerful because it provides both **global explanations** (by aggregating Shapley values across many predictions to see overall feature importance) and **local explanations** for each individual prediction. Many practitioners favor SHAP for its consistency and theoretically sound foundation, and tools exist to visualize SHAP values with summary plots, force plots, and more for easy interpretation.

Aside from LIME and SHAP, there are numerous other post-hoc techniques: **saliency maps** in computer vision highlight which pixels in an image influenced a classification (useful for explaining why an image was labeled a certain way), **counterfactual explanations** pose “what-if” scenarios (e.g., *“If the applicant had a slightly higher income, the model would have approved the loan”*), and **concept-based methods** try to explain decisions in terms of high-level concepts rather than raw features (especially in image and text domains). There are also toolkits like IBM’s AI Explainability 360 (AIX360) and Microsoft’s InterpretML that bundle multiple algorithms and provide a unified interface for generating explanations.

It’s important to note that explainability techniques can be used in combination

. For example, a team might use an interpretable model for the core decision and then a SHAP analysis on top for additional nuance. Or they might use global feature importance to identify potential issues, then drill down with local explanations on specific cases. In practice, the choice of explainability method depends on the audience and requirements: **Executives or end-users** may prefer simple natural-language or visual explanations (even if approximate), whereas **data scientists** might inspect detailed weight vectors or decision rules to audit a model. The good news is that the ecosystem of XAI tools is growing, making it easier to attach an “explanation layer” to just about any AI pipeline. The key is to ensure the explanations themselves are understandable and appropriately accurate. A good explainability solution should offer **human-friendly outputs** (avoid unnecessary technical jargon or complexity), provide both **local and global perspectives** on model behavior, maintain **traceability** (so that one can document how decisions are made and replay the reasoning later if needed), and ideally integrate into existing workflows (for instance, through dashboards or reports that decision-makers can easily use).

Finally, it’s worth mentioning a trade-off that often arises: sometimes, to make a model more explainable, one might sacrifice a bit of accuracy or complexity. Simpler models are easier to explain but might not capture all patterns; more complex models capture more but are harder to interpret. Finding the right balance – or using explainability techniques that minimize loss of accuracy – is part of the art of deploying AI responsibly. Encouragingly, research and best practices are continually improving, so organizations no longer have to choose between a **“powerful model” and a “transparent model”** – with modern XAI methods, you can often have both to a satisfying degree.

***

## Inference and Explanation Integration in Production

<figure><img src="/files/3beMGesEa6IvvwIUs0Nu" alt=""><figcaption></figcaption></figure>

***

## Mechanistic Interpretability vs. Explainability

Amid discussions of explainable AI, you might also hear the term **“**[**mechanistic interpretability**](https://www.transformer-circuits.pub/2022/mech-interp-essay)**.”** While closely related to explainability, mechanistic interpretability has a more specific, technical focus. In simple terms, mechanistic interpretability is the **study of \_reverse-engineering**\_\*\* a trained AI model to understand exactly how its internal parts operate\*\*. Scholars use this term especially in the context of complex neural networks (like the large deep learning models powering today’s language AI). The idea is analogous to popping open the hood of a car engine: mechanistic interpretability tries to dissect an AI model’s “gears and circuits” – its neurons, layers, and weights – to figure out what each component is doing and how they collectively implement the model’s function.

Traditional explainability approaches (like LIME or SHAP discussed above) tend to treat the model as a black box and focus on explaining inputs and outputs – they tell you *what* inputs led to *what* outputs. Mechanistic interpretability, by contrast, wants to **understand the \_how**\_\*\* at a structural level\*\*. For example, in a large language model like GPT, researchers might try to identify individual neurons or groups of neurons that correspond to interpretable concepts (like a neuron that activates for names of foods, or a cluster of neurons that tracks grammar structure). They might analyze how information flows through the network’s layers or how internal representations transform as the model processes an input. This field has seen fascinating progress in recent AI research. Pioneering work by teams looking at networks like GPT-2 has revealed the presence of “circuits” – combinations of neurons and weights that together perform a semantic task (for instance, one circuit might be responsible for a model’s ability to match opening and closing parentheses in text). By **reverse-engineering these neural circuits**, researchers inch closer to a mechanistic understanding of why the model outputs what it does.

Why does mechanistic interpretability matter? One motivation is **AI safety and reliability**. If we can deeply understand a model’s internal mechanics, we might detect flaws or emergent undesirable behaviors (like a tendency to produce biased outputs or to “trick” its objective function in unintended ways) before they cause harm. It can also guide us in correcting or editing models – for example, if a specific neuron consistently triggers toxic language in a generative model, a mechanistic insight might allow us to modify or constrain that part of the network. In essence, mechanistic interpretability is pushing beyond treating the model as a black box that we explain externally; it seeks to **open the black box** and read the model’s “source code” that it unwittingly wrote during training. This is an active frontier: it’s **young research**, and currently feasible mostly for smaller-scale networks or specific components. No one can yet fully interpret the likes of GPT-4 or other extremely large models – the complexity is staggering – but the work has begun in earnest. Over time, advances in this area might complement higher-level explainability techniques, giving us a multi-layered understanding of AI: from the detailed circuits up to the user-facing explanations.

To put it succinctly, **explainability** (in the XAI sense) typically focuses on providing useful *external* explanations for humans (often answering “why did the AI make X decision?” in user-friendly terms), whereas **mechanistic interpretability** aims to *internally* understand the AI’s actual mechanics (“how is the computation implemented in the network?”). Both are important. For most organizations today, the priority is explainability in the XAI sense – delivering immediate, practical insights and justifications for AI decisions. Mechanistic interpretability is more of a research endeavor that could, in the long run, make AI systems more transparent from the ground up. One can imagine a future where, thanks to mechanistic insights, our AI models are built in ways that are inherently easier to interpret. Until then, XAI techniques provide the bridge that allows humans to trust and supervise the AI we have now.

## Explainability in the Data Analytics Ecosystem

How does AI explainability fit into the day-to-day world of data teams and analytics workflows? In modern data-driven organizations, AI models are not standalone curiosities – they are woven into a broader **analytics ecosystem** that includes data warehouses, business intelligence (BI) tools, data pipelines, and decision-making processes. Explainability serves as a crucial link in this ecosystem, ensuring that the insights derived from AI are **accessible and actionable to the people who need them**.

Consider a typical scenario in a data-driven company: a machine learning model might be built to predict something like customer churn, sales forecasts, or risk scores. The predictions from this model could feed into a BI dashboard seen by a marketing manager, or into an automated system that triggers actions (like reaching out to at-risk customers). If that model is a black box, the managers and analysts downstream are left in the dark about *why* the numbers are what they are. This is where explainability comes in. By integrating explainable AI, those dashboards or reports can display not just the **“what” (the prediction)** but also a digestible version of the **“why.”** For instance, a churn risk dashboard might list the top three factors contributing to each customer’s risk score (e.g., *“low engagement in last 30 days”*, *“reported an issue with product quality”*). This transforms AI from a mysterious oracle into a collaborative tool – the data team and business team can have a conversation around the model’s findings, grounded in the model’s reasoning.

New analytics tools are emerging that embody this principle. Take [Dot, the AI data analyst](https://getdot.ai), as an example. Dot is essentially a conversational interface that lets users ask questions of their data in plain English and get answers backed by AI. For such an AI assistant to be useful in a business setting, it needs to not only fetch numbers or make predictions, but also to **explain and contextualize those results**. If a user asks, “Why did our revenue dip last quarter?”, an AI like Dot might analyze the data and respond with a narrative insight (e.g., *“Revenue fell 5% due to lower sales in Region X and an increase in product returns; the model indicates the biggest factors were a supply shortage and a drop in repeat customers in that region.”*). The value of this kind of tool is that it provides *instant, actionable insight* – and it’s the explainability component (highlighting key drivers and context) that makes the insight actionable and trustworthy, rather than a black-box answer. By linking to the underlying data and reasons, explainable AI assistants ensure that data **democratization** doesn’t come at the cost of rigor or trust. Everyone from a data engineer to a business stakeholder can understand the “story” behind the data, thanks to the explainability built into the AI analytics workflow.

On the engineering side, explainability is also becoming part of the machine learning operations stack. We see model monitoring services and data platforms incorporating explainability features. For instance, major cloud and data warehouse platforms have started to provide **built-in explainability** for models deployed on their infrastructure. A recent development in Snowflake (a popular cloud data platform) is a feature that allows users to compute **Shapley values for models directly within the data warehouse**, so data scientists can easily examine feature contributions without exporting data to separate tools. This kind of integration means that explainability is not an afterthought but a **native part of model deployment**: whenever a prediction is made, an explanation can be logged or served up as well. It also addresses practical concerns like data governance and security – by doing explainability in-platform, sensitive data doesn’t have to be shuffled around to third-party services for analysis.

Explainability also complements **data governance and cataloging** efforts. Tools from companies like Alation, Collibra, or Atlan help organizations keep track of their data assets, data lineage, and ensure data quality. When models are producing insights that feed into critical decisions, treating those models and their outputs as governed assets is important. Explainability reports (like which factors influenced a decision, or whether the model is behaving within expected bounds) can be logged as part of governance records. This creates an audit trail for automated decisions, similar to how we maintain logs for traditional business processes. In regulated industries, such an audit trail is invaluable for demonstrating compliance. Even in more agile environments, it’s useful for **knowledge sharing**: future team members can understand past model decisions by reviewing explanations, much like a scientist keeps a lab notebook of experimental results and interpretations.

In summary, AI explainability fits into the data analytics ecosystem as the **translation layer** between complex models and human decision-makers. It ensures that AI-driven insights are not siloed with data scientists but are **shared in an understandable form across the organization**. By embedding explainability into data tools, from AI assistants like [Dot](https://getdot.ai) to enterprise data platforms, organizations enable a more collaborative and transparent use of AI. The result is that insights generated by AI can be trusted and acted upon, accelerating the data-driven decision culture that so many organizations strive for. Explainability, in this sense, amplifies the value of AI by linking it tightly with human context and judgment in the analytics value chain.

## Use Cases and Applications of Explainable AI

Explainable AI is not just a theoretical nice-to-have – it’s being applied in a wide array of industries and scenarios where understanding AI decisions is mission-critical. Let’s explore a few representative use cases to see how explainability adds value:

* **Financial Services (Credit and Lending):** Banks and fintech companies use AI models to assess credit risk and decide whether to approve loans or credit cards. These models must comply with regulations that often require giving customers an explanation if they are denied credit. An explainable AI model in this context might produce a credit score *and* a list of key factors (e.g., high credit utilization, short credit history) that led to that score. Not only does this fulfill regulatory requirements, but it also helps loan officers review and trust the model’s judgement. Moreover, by examining common explanation patterns, a bank might identify if its model is inadvertently using proxies for protected characteristics (like race or gender), enabling it to address potential biases proactively. Explainability thus supports **fair lending practices** and helps maintain transparency with consumers. In the realm of algorithmic trading or asset management, explainability is used internally to ensure models aren’t taking on hidden risks – for example, a trading model might be required to explain which market signals or indicators are driving its decisions, so that human analysts can double-check that those align with sound strategy (and not, say, picking up a transient anomaly).
* **Healthcare:** AI is increasingly used for diagnosing diseases from medical images, recommending treatments, or predicting patient outcomes. Doctors are rightly cautious about using such tools unless they can **justify the reasoning**. For instance, if an AI model analyzes an X-ray and flags a potential tumor, an explainable system could highlight the specific area of the image and the features that led to that conclusion (perhaps texture patterns or shapes that the model associates with malignancy). This acts like a second set of eyes for the radiologist – one that can point and say *“look here, this tissue looks irregular in a way similar to past cancer cases.”* In predictive health (like models that forecast which patients are at risk of complications), an explanation might be a simple list: *“Key factors: age, blood pressure trend, and a specific lab result were the top contributors to this risk prediction.”* Such transparency is crucial for clinicians to trust the AI and incorporate its findings into their decision-making. It also enables **patient-facing explanations** – a doctor can better communicate to a patient why an AI-influenced diagnosis was made, improving patient understanding and trust in the overall care process. More broadly, explainability in healthcare AI supports compliance with medical accountability standards and can accelerate the adoption of AI by building a bridge between data science and clinical expertise.
* **Insurance:** Insurance firms use AI for underwriting (deciding policy terms or pricing based on risk) and claims processing (detecting fraud or estimating damages). In underwriting, explainable AI can clarify why a certain applicant is deemed higher risk – for example, *“The model increased the auto insurance premium due to the applicant’s young age and recent accident history, which statistically elevate risk.”* This transparency is important both internally (actuaries and risk officers need to ensure the model’s decisions make business sense and are not unfairly discriminatory) and externally (customers may inquire why their quoted rate is high; an explanation fosters clarity and trust). In claims, if an AI flags a claim as potentially fraudulent, an explanation is essential for the investigation team to follow up – perhaps the AI noticed that *“the claimed incident description is eerily similar to 10 past fraudulent cases”* or *“the claimant’s vehicle location data doesn’t match the accident report.”* These clues guide human investigators. Notably, regulators like state insurance commissions are starting to demand such explanations to ensure AI-driven insurance decisions are **transparent and equitable**. As mentioned earlier, California now requires that if an algorithm is used to make an adverse insurance decision, the company must be able to explain it. This trend is likely to expand to other jurisdictions and lines of insurance.
* **Manufacturing and IoT (Predictive Maintenance):** In industrial settings, AI models predict equipment failures or quality issues on the production line. Explaining these predictions can save a lot of time and money. If a model predicts that a certain machine is likely to fail in the next week, the maintenance team needs to know why. An explainable system might indicate that *“vibration sensor readings on the motor have been fluctuating beyond normal range and temperature is rising”*, pointing technicians to the root cause (perhaps a misaligned shaft or impending bearing failure). Without that explanation, the team might not trust the alert or might not know where to begin looking. Similarly, in quality control, if an AI flags a batch of products as defective, an explanation could highlight which sensor readings or process conditions contributed, enabling engineers to pinpoint the issue (like a particular valve that was set incorrectly). Here, explainability ensures that AI-driven insights are **practical** – they translate into concrete operational actions and empower engineers with diagnostic understanding rather than just alarm bells.
* **Retail and Marketing:** Retailers use AI for personalized recommendations, pricing optimization, and churn prediction. While recommending a product or personalizing a price might not seem like a life-or-death decision, explainability still adds value. It can help marketers and product teams understand customer segments better. For example, an AI might personalize an e-commerce homepage differently for a user, and an explanation system might reveal *“Customer is shown more sportswear because their browsing and purchase history indicates a strong preference for fitness-related items.”* This is similar to how streaming services sometimes tell you “Because you watched X, we recommend Y.” These friendly explanations increase user engagement and trust – customers feel the system “knows” their preferences rather than randomly tossing items at them. Internally, if a promotion model suggests a discount for certain customers, analysts will want to know the rationale (perhaps those customers have a high estimated lifetime value but haven’t purchased recently, etc.). By explaining model-driven marketing actions, companies can refine their strategies and ensure they align with customer relationships (for instance, not offering a deep discount to someone who likely would pay full price, which an explanation might catch by highlighting factors like price sensitivity scores). In customer churn models (predicting who will stop using a service), as we discussed, explanations guide retention efforts – e.g., *“This subscriber is at risk of churning due to low usage and a recent price hike; proactive outreach with a special offer might retain them.”*
* **Public Sector (Justice and Security):** Government agencies are experimenting with AI for things like allocating social services, flagging tax fraud, or even aiding judicial decisions (e.g., risk assessment scores in courts). These are extremely sensitive areas where transparency is paramount. If an AI system recommends increased screening for certain tax returns, it must explain the red flags (unusual deduction patterns, income inconsistencies, etc.) so that auditors can double-check and citizens can be treated fairly. In criminal justice, any algorithm used to assess, say, the likelihood of re-offense (recidivism risk scores) has faced criticism when it’s a black box. Explainability would require such a system to spell out the factors (prior offenses, age, etc.) and how they combine into the recommendation, allowing a judge or parole board to weigh that alongside other context. In practice, due to the controversy, some jurisdictions have rolled back on using black-box scoring entirely; but if such tools are to be considered, they likely must come with **transparent explanations** to be publicly and legally acceptable.

Across all these use cases, a common theme is that explainability aligns AI with **human values and judgment**. By illuminating the AI’s reasoning, domain experts and decision-makers can critique and adjust the AI’s output if needed. This **human-AI collaboration** is often the ideal scenario: the AI provides speed, scale, and pattern-recognition, while the human provides oversight, ethical judgment, and contextual understanding. Explainability is what facilitates this partnership. It’s also worth noting that explainability can sometimes reveal when an AI is actually **wrong or overconfident**.

For example, if an explanation for a medical diagnosis doesn’t make medical sense, that’s a red flag to the doctor that the model might have erred for a strange reason (maybe a quirk in the training data). Thus, explanations not only tell us when to trust the model, but also when *not* to trust it.

In summary, explainable AI is applied wherever AI decisions intersect with real-world decisions that matter – which is an ever-growing set of domains. From finance to healthcare to everyday business analytics, explainability ensures AI’s outputs are interpretable and actionable. It transforms AI from a mystical oracle into a well-lit instrument panel that humans can read and navigate by.

## Choosing and Using Explainability Tools: What to Look Out For

If you’re considering implementing explainability in your AI projects or evaluating an **XAI solution** to purchase, there are several key factors and potential pitfalls to keep in mind. Not all explainability tools are created equal, and using them correctly is as important as the tool itself. Here are some guidelines on what to look out for:

**1. Clarity and Human-Friendliness:** The entire point of explainability is to make AI understandable to humans, so the explanations should be presented in a clear, intuitive manner. When evaluating a tool, check the format of its outputs. Do they produce textual explanations, visualizations, or charts that a non-data-scientist can grasp? For example, a tool might output: *“Feature X contributed +0.5 to the prediction, Feature Y contributed -0.2”* alongside a simple bar chart. That might be perfectly clear to an analyst. But if another tool spits out a complex decision tree or a large table of numbers, it may defeat the purpose if the audience can’t easily interpret it. **Human-readable explanations** are a must. Also, consider if the tool supports natural-language summaries or interactive exploration, which can enhance understanding for different users.

**2. Local vs. Global Explanations:** Different stakeholders have different needs – sometimes you need to explain a specific decision (local), and other times you need to understand the model’s overall logic (global). A good explainability solution should ideally provide both **local interpretability (per-instance explanations)** and **global interpretability (an overall view of model behavior)**. For example, local explanations help answer, “Why did the model reject *this* loan applicant?” whereas global explanations answer, “In general, what factors drive the model’s loan decisions?” When selecting a tool, see if it can drill down into individual predictions as well as offer aggregate insights. Some tools specialize in one or the other, so your choice might depend on whether your primary use case is case-by-case explanations (important in customer-facing or audit scenarios) or broader model understanding (important for model validation and regulatory documentation).

**3. Compatibility and Scope:** Ensure the explainability method supports the types of models and data you are using. Some techniques are **model-agnostic** (they can explain any model by treating it as a black box, like LIME and SHAP), while others are **model-specific** (for example, methods that only work for tree-based models or only for neural networks). If you have a mix of model types, a unified tool might simplify your workflow. Also consider data modality: do you need to explain image or text models, or just tabular data? Techniques like saliency maps are image-specific, while others like LIME have variants for text vs. tabular. Make sure the solution covers your scope. It’s wise to test it on a known scenario – say, train a small model and see if the explanations align with what you expect – before rolling out broadly.

**4. Performance and Scalability:** Generating explanations can sometimes be computationally intensive. For instance, SHAP values are powerful but can be slow to compute for large datasets or very complex models, because they involve evaluating many combinations of features. When integrating explainability into real-time systems, consider the performance implications. Some tools provide approximate or faster versions of their algorithms (e.g., using sampling to approximate SHAP values, or caching results). If you’re buying a solution, ask about its **scalability** – can it handle explaining thousands of predictions in a batch? Does it introduce latency if you want an explanation on the fly for a single prediction? In high-frequency decision environments (like algorithmic trading or fraud detection with milliseconds to respond), a heavy explainability process may not be feasible in the loop, so you might use it offline for analysis rather than inline for every decision. Balance the need for thorough explanation with the practical performance needs of your use case.

**5. Integration with Workflow:** The best explainability tool is one that can be easily woven into your existing processes. Look for tools that have **APIs or interfaces** that plug into your infrastructure – for example, can it integrate with your Python notebooks, your data warehouse, or your BI platform? If you use Snowflake, the fact that Snowflake now supports certain explainability functions natively might tilt you toward using those for convenience. If your team is already using a tool like SageMaker or Databricks for model development, check if they offer built-in explainability features or if third-party libraries (like SHAP, AIX360, etc.) can be added. Also, consider how the explanations will reach end-users or decision-makers: do you need to export them to a dashboard, or generate PDF reports, or send alerts? A smooth integration means explanations will actually reach the people who need them, in the tools they already use.

**6. Accuracy and Faithfulness of Explanations:** This point is subtle but crucial. An explanation is only useful if it accurately reflects the model’s true reasoning. Some explainability methods produce very convincing-sounding explanations that are *approximations* of the model behavior, and if used blindly they could be misleading. For example, a simple linear explanation from LIME might suggest a certain relationship, but it’s only true locally and might not hold globally. There have been cases in research where certain explanation tools could be “fooled” by adversarial examples – meaning the tool gives the same explanation for two very different inputs, or misses a feature interaction. While most mature tools are robust for normal use, **users should not place blind trust in any single explanation**. It’s wise to sanity-check explanations (do they make sense domain-wise?) and use multiple methods if something looks off. When evaluating a solution, ask if it provides any confidence metrics for its explanations or if it allows manual inspection. For instance, some tools might highlight potential uncertainty (e.g., “this explanation is less reliable because the instance is out-of-distribution”). Additionally, be aware of the **limitations of post-hoc explanations** – they deduce influences based on output behavior and might not capture complex interactions perfectly. This is an ongoing area of research, so staying informed via validation tests is key. In regulated settings, you might even need to validate the explanation method itself (proving that the explanations are faithful to the model and not just plausible-sounding stories).

**7. Traceability and Audit Trails:** Especially for enterprise use, consider how explanations will be stored and reviewed. If an AI decision is later challenged (by a customer or regulator), having a recorded explanation for that specific decision can be extremely helpful. Some explainability solutions log every explanation generated along with timestamps and model versions – effectively creating an **audit trail of model decisions**. This is part of a larger AI governance practice. If you’re implementing your own solution, you may want to build in such logging. Traceability also means documenting the version of the model and data that the explanation pertains to (since models can be updated). When comparing tools, see if they support linking explanations with model versioning or if they offer a dashboard to review past explanations. This feature might not be critical for an experimental project, but for production systems it can save a lot of headaches later.

**8. Ethical and Bias Considerations:** A good explainability tool can also become a lens to examine fairness and bias in your model. Features that consistently appear in explanations can hint at potential biases. For instance, if “zip code” shows up frequently as a top factor in a credit model’s explanations, and you know zip code correlates with race in your region, that’s a flag to investigate fairness. Some specialized tools incorporate fairness metrics alongside explanations, or allow “counterfactual” analysis (e.g., “if we change this sensitive attribute, does the decision change?”). While this goes a bit beyond core explainability, it’s something to look for if your organization has strong **Responsible AI** requirements. At minimum, ensure that the explainability solution doesn’t hide or obscure such factors. There have been examples where an AI vendor, attempting to avoid controversy, might try to suppress certain kinds of information in explanations – but it’s usually better to know and address biases than to keep them hidden. Transparency is a double-edged sword: it can reveal uncomfortable truths about your model. Be prepared to act on what you learn (e.g., retrain the model, add constraints, or improve data) rather than assuming the explanation tool will magically solve bias.

In conclusion, when using or buying an explainability solution, **do your due diligence**: check that the features align with your needs (interpretability, traceability, integration), verify the tool’s output on known cases, and remain critical of the results. Explainable AI is a powerful aid, but it’s not infallible. The goal is to choose a tool that genuinely helps humans make sense of AI in your specific context. A thoughtful selection and implementation will yield an explainability process that enhances trust in your models and leads to better outcomes. On the other hand, neglecting these considerations could result in confusion or false confidence. So, treat explainability as you would any analytical process – with rigor and care – and it will greatly enrich your AI initiatives.

## The Future of AI Explainability

As AI systems continue to advance and permeate every industry, the importance of explainability will only grow – and so will the techniques to achieve it. Looking ahead, several trends indicate where AI explainability is headed and how it will shape the **future of AI and analytics**:

**1. From Optional to Mandatory:** What is today considered best practice could soon become a baseline requirement. We are already seeing regulatory shifts that make explainability a built-in expectation for AI systems. The European Union’s upcoming AI regulations, for example, are poised to enforce transparency for “high-risk” AI applications. Other countries and states are exploring similar rules. In the near future, it’s likely that any AI system impacting consumer rights (credit, employment, healthcare, etc.) will legally need to provide explanations for its outputs. Even beyond legal mandates, public opinion and business ethics are trending toward demanding more transparency. Companies that cannot explain their AI’s decisions may find themselves at a competitive disadvantage or under public scrutiny. In contrast, those who embrace explainability can earn **digital trust** from their customers and partners – which, as studies suggest, correlates with better financial performance. In short, explainability is moving from a niche concern of data scientists to an organization-wide value, akin to data security or privacy. We can expect future AI development frameworks to include explainability as a first-class component, with documentation and user-interface considerations baked in.

**2. Advances in Explainability Techniques:** The research community is actively working on making explanations more informative, more precise, and applicable to cutting-edge AI models. One challenge on the horizon is explaining the new wave of **large language models (LLMs)** and other “foundation models” (like giant image or multimodal models). These models, like GPT-3 or GPT-4 and their successors, are extremely complex – but they are also being deployed widely (in chatbots, content generation, etc.). Traditional tools like LIME or SHAP may not scale neatly to such massive models or may not capture their sequential decision processes. In response, we’re likely to see new kinds of explainers that are tailored to these models. For instance, researchers are exploring how to get language models to **explain themselves** by generating rationales for their answers. Imagine asking a future AI, “Should we approve this loan?” and it not only says “Yes” or “No,” but also gives a coherent explanation: *“No, because the applicant’s debt-to-income ratio is above our threshold and they have a recent delinquency. Historically, those factors led to defaults.”* In some initial experiments, prompt-based techniques allow language models to produce such explanations (though ensuring their correctness is an ongoing challenge). We may also see more use of **counterfactual explanations** (“if X were different, the decision would be different”) to complement traditional feature-attribution methods, as counterfactuals can be very intuitive for users. Additionally, tools will improve to handle more complex data – like explaining a model that takes in an entire **time series** or **graph** data. Visualization techniques are bound to become more interactive and user-friendly, possibly leveraging VR/AR for very complex scenarios, though that’s further out.

**3. Mechanistic Interpretability Achievements:** The earlier discussion on mechanistic interpretability hints at a future where we might actually open the black box in a more fundamental way. While still largely in the research domain, there’s optimism that progress here will translate into practical benefits. For example, if researchers succeed in **reverse-engineering significant portions of a model’s cognition**, future AI systems might come with a map of their neural circuits that developers can inspect. It could become feasible to debug a neural network almost like debugging software, identifying which “subroutine” (set of neurons) caused an undesired output. This could revolutionize how we trust AI – you wouldn’t have to take the entire model’s output on faith if you can pinpoint the exact component responsible for a behavior. In an optimistic scenario, mechanistic interpretability could also help with **model editing** – surgically fixing parts of a model without retraining from scratch. That would be a game-changer for maintaining large AI systems. In the shorter term, insights from mechanistic studies are informing simpler explainability tools. For example, knowing that certain neurons represent certain concepts can lead to more concept-driven explanations (e.g., “the model thinks this text is positive because it detected the concept of ‘praise’ and ‘achievement’ in the writing”).

**4. Integration with AI Governance and Automated Monitoring:** We will likely see explainability tightly integrated with automated monitoring systems. Instead of manually generating explanations when something goes wrong, future AI ops platforms might continuously monitor not just model accuracy but also model **explainability profiles**. For instance, if an AI in production suddenly starts making decisions based on an unusual factor (say a credit model suddenly gives high importance to an applicant’s phone area code, which wasn’t a major factor before), the system could trigger an alert. This is akin to anomaly detection but on the explanation level. Such automated explainability monitoring could catch issues like data drift or bias drift early. It could also enforce constraints – like if there’s a policy that certain features should not significantly influence decisions, the monitor can ensure explanations stay within those bounds (and flag if not). In practice, this means explainability will be part of the continuous delivery and monitoring cycle of ML: just as tests and metrics are used to decide if a model can be deployed or needs retraining, explanation patterns will be analyzed for compliance and reasonableness as a standard procedure.

**5. User-Centered Explanations:** The future will also bring more focus on *who* the explanation is for. A one-size-fits-all explanation might not be ideal; instead, AI systems may offer different layers of explanation for different users. An executive might get a one-sentence summary explanation, an operational manager might get a detailed breakdown with visuals, and a data scientist might get a technical report. We see early moves in this direction: some AI platforms allow the customization of explanation content. There’s research on **explanation interfaces** that adapt to a user’s level of expertise (for example, a doctor might get medical terminology in the explanation, whereas a patient gets layman’s terms for the same model output). Making explanations more **interactive** is another direction – letting users ask follow-up questions about a decision (“What if the income were higher?” or “How much did factor X influence this prediction compared to factor Y?”) and having the system answer in real time. With natural language capabilities of AI improving, this kind of dialogue about the AI’s reasoning could become a standard feature, effectively allowing users to “interrogate” the model as they would a human decision-maker. This has huge implications for acceptability – an AI that can engage in a two-way explanation conversation might overcome a lot of the skepticism people have when they can’t get a straight answer out of a machine.

**6. Cultural and Organizational Change:** Lastly, the spread of explainable AI will likely change organizational culture around AI. When AI decisions are transparent, it encourages accountability and cross-disciplinary involvement. We might see **AI ethics committees** and **governance boards** within organizations routinely reviewing explanation reports, much like financial audit committees review audit statements. Explainability could thus shift AI development from being purely the domain of technical teams to a more collaborative process with input from compliance officers, domain experts, and even representatives of those affected by the AI’s decisions. This is a positive direction – it embeds AI into the fabric of business processes with appropriate oversight. Tools will probably evolve to serve these committees – think dashboards that show how each model in production is making decisions, with drill-down capabilities, scenario analysis, etc., all in a user-friendly way. In a sense, AI explainability might become **part of the KPI framework** for AI initiatives: companies could track not just what their AI did, but how it did it, and use that knowledge to improve both the AI and the business.

In conclusion, the future of AI explainability is poised to make AI systems even more **transparent, interactive, and aligned with human needs** than they are today. We can expect explainability to be deeply ingrained in AI tools and platforms – a default rather than an add-on. As that happens, the relationship between humans and AI will become more like a partnership of colleagues than a mysterious master-servant dynamic. When AI can answer the question “why did you do that?” as comfortably as a person can, we will truly be in an era of AI that is **accountable and trustworthy by design**. Achieving that at scale is no small task, but the trends suggest we are well on our way. After all, the ultimate promise of explainable AI is that it enables us to harness the power of advanced algorithms **without losing visibility or control**, ensuring that these systems remain beneficial, fair, and aligned with our goals. Organizations that invest in this capability now are not just solving a technical problem – they are laying the groundwork for a future in which AI is an open book and a collaborative ally in every sense of the word.


# Artificial Intelligence

### What Exactly is Artificial Intelligence?

Artificial intelligence represents a fundamental shift in how machines can process information and make decisions. At its core, AI refers to computer systems that can perform tasks typically requiring human intelligence—reasoning, learning, problem-solving, perception, and [language understanding](https://www.iso.org/artificial-intelligence). The term itself was coined by Stanford professor John McCarthy in 1955, five years after Alan Turing proposed his famous test for [machine intelligence](http://jmc.stanford.edu/general/index.html).

Rather than following pre-programmed instructions like traditional software, AI systems learn from data and improve their performance over time. They analyze patterns, draw conclusions, and [make predictions or recommendations](https://www.atlassian.com/blog/artificial-intelligence/artificial-intelligence-101-the-basics-of-ai) based on what they've learned. Think of AI as teaching a computer to recognize patterns the way a child learns to identify objects—through repeated exposure and feedback, gradually building understanding and capability.

<figure><img src="/files/93NZy8IGlFsHSkP7pNLX" alt=""><figcaption></figcaption></figure>

*This diagram illustrates the core AI learning cycle, showing how systems continuously improve through feedback and iteration.*

The distinction between AI and traditional computing lies in adaptability and autonomous learning. While conventional programs execute specific instructions, AI systems can generalize knowledge, adapt to new situations, and improve performance without explicit reprogramming. This capability enables AI to tackle complex, nuanced problems that would be impossible to solve through traditional programming approaches.

### Why Do Organizations Need Artificial Intelligence?

Organizations today face an unprecedented challenge: exponential data growth coupled with the need for faster, more accurate decision-making. Traditional approaches to data analysis and business intelligence, while valuable, often fall short in handling the volume, velocity, and variety of modern [data streams](https://www.teradata.com/insights/ai-and-machine-learning/ai-for-data-analysis-and-visualization). This is where AI becomes not just beneficial, but essential for competitive survival.

#### The Business Intelligence Limitation

Traditional business intelligence platforms excel at historical reporting and descriptive analytics. They gather data from multiple sources, clean it, and present it through [dashboards and visualizations](https://www.qlik.com/us/business-intelligence-platform). However, these systems primarily answer "what happened" rather than "what will happen" or "what should we do about it". They require human analysts to interpret patterns, draw conclusions, and make recommendations—a process that's time-consuming and limited by human cognitive capacity.

#### The AI Advantage in Modern Business

AI transforms this paradigm by automating complex analytical tasks and providing [predictive and prescriptive insights](https://dev.to/archit_nandan_ff/the-rise-of-ai-powered-analytics-the-transformative-power-of-ai-powered-analytics-4824). Organizations implementing AI-driven analytics report significant improvements in decision-making speed and accuracy. According to recent research, businesses using AI in analytics are five times more likely to make faster decisions than those relying solely on traditional methods.

The value proposition extends across multiple dimensions:

**Predictive Capabilities**: AI systems can forecast future trends, customer behaviors, and potential risks with remarkable accuracy. Retailers use AI to predict demand patterns, enabling optimal inventory management and [reducing waste by up to 25%](https://www.consultancy-me.com/news/10043/six-use-cases-for-artificial-intelligence-ai-in-business). Financial institutions leverage AI for credit risk assessment, achieving 90% accuracy in default predictions.

**Real-Time Processing**: Unlike traditional BI systems that work with historical data, AI can [process and analyze data streams in real-time](https://www.oracle.com/be/artificial-intelligence/artificial-intelligence-analytics/). This capability is crucial for industries where immediate responses matter—from fraud detection in financial services to predictive maintenance in manufacturing.

**Automated Insight Generation**: AI systems can automatically identify anomalies, surface hidden patterns, and generate actionable recommendations without human intervention. This automation frees data analysts to focus on strategic initiatives rather than routine data processing tasks.

**Scale and Complexity Management**: AI excels at processing vast amounts of unstructured data from diverse sources—social media sentiment, customer reviews, sensor data, and more. Traditional BI systems struggle with this variety and volume of information.

#### Transforming Decision-Making Processes

The most significant organizational benefit of AI lies in its ability to democratize [data-driven decision-making](https://getdot.ai). Tools like Dot, the AI data analyst, enable non-technical users to ask complex business questions in natural language and receive instant, accurate [insights](https://dang.ai/tool/ai-data-analysis-getdot-ai). This accessibility means that decisions can be made at the point of need, rather than waiting for analysts to generate reports.

Consider a marketing manager who can ask "Why did our conversion rates drop in the Northeast region last month?" and receive not just the answer, but the underlying factors and recommended actions. This level of accessibility transforms how organizations operate, making them more agile and responsive to [market changes](https://blog.getdot.ai/introducing-deep-analysis-the-next-generation-of-ai-powered-analytics-agents-476d6ca03e86).

#### The Economic Impact

The economic implications are substantial. AI is projected to add [$4.4 trillion to the global economy annually](https://explodingtopics.com/blog/future-of-ai). Organizations that embrace AI-driven analytics report measurable improvements in key performance indicators, reduced operational costs, and enhanced [customer satisfaction](https://imaginovation.net/blog/al-analytics-businesses/). The technology pays for itself through improved efficiency, reduced errors, and better strategic decision-making.

#### Beyond Traditional Analytics

AI enables organizations to move beyond reactive reporting to proactive strategy. Predictive analytics help businesses anticipate market shifts, customer needs, and operational challenges before they become problems. Prescriptive analytics go further, recommending specific actions to achieve desired outcomes.

This shift from "what happened" to "what should we do" represents a fundamental change in how organizations operate. Companies can now optimize marketing campaigns in real-time, adjust pricing strategies based on market conditions, and proactively address customer concerns before they [escalate](https://blog.hubspot.com/sales/ai-business-analytics).

### Different Types of Artificial Intelligence

Understanding AI's various forms helps organizations choose the right approach for their specific needs. AI systems can be categorized by their capabilities and learning methods, each serving different business purposes.

#### AI by Capability Levels

**Narrow AI (Weak AI)** represents the current state of artificial intelligence technology. These systems excel at specific, well-defined tasks but cannot transfer knowledge to other domains. Examples include voice assistants like Siri recognizing speech, recommendation engines suggesting products, and fraud detection systems identifying suspicious [transactions](https://bernardmarr.com/the-10-best-examples-of-how-companies-use-artificial-intelligence-in-practice/). While "narrow," these systems can outperform humans in their specialized areas.

**General AI (Strong AI)** remains theoretical—systems that could match human intelligence across all domains. These would possess human-like reasoning, creativity, and adaptability, capable of learning any task that humans can perform. Current AI systems, regardless of their sophistication, fall into the narrow category.

**Super AI** represents hypothetical systems that would exceed human intelligence in all areas. This remains in the realm of speculation and long-term research rather than practical application.

#### AI by Learning Approach

**Supervised Learning** systems learn from labeled examples, like training a system to recognize spam emails by showing it thousands of messages already classified as spam or legitimate. Most business AI applications use supervised learning for tasks like customer segmentation, demand forecasting, and quality control.

**Unsupervised Learning** finds patterns in data without predetermined labels, discovering hidden structures or relationships. Businesses use this for customer behavior analysis, market segmentation, and anomaly detection where the patterns aren't known in advance.

**Reinforcement Learning** systems learn through trial and error, receiving rewards or penalties for their actions. This approach powers autonomous systems, game-playing AI, and optimization algorithms for complex [business processes](https://www.icaew.com/insights/viewpoints-on-the-news/2024/nov-2024/types-of-ai-how-are-they-classified).

#### Functional AI Categories

**Reactive Machines** perform specific tasks without memory or learning capability. IBM's Deep Blue chess computer exemplifies this type—excellent at chess but unable to apply that knowledge elsewhere.

**Limited Memory AI** uses historical data to inform current decisions. Most modern business AI falls into this category, analyzing past patterns to predict future outcomes or recommend actions.

**Theory of Mind AI** would understand that others have beliefs, desires, and intentions—currently under development. Future customer service AI might possess this capability, better understanding customer emotions and motivations.

**Self-Aware AI** would possess consciousness and self-understanding—purely theoretical at present.

For business applications, the focus remains on narrow AI with supervised and unsupervised learning capabilities. These systems provide immediate, measurable value while remaining controllable and interpretable.

### What's a Business Intelligence Platform vs. Artificial Intelligence?

The distinction between business intelligence platforms and artificial intelligence reflects a fundamental difference in approach to data analysis and [decision-making](https://valueworks.ai/understanding-business-intelligence-platforms-benefits-and-features/). While both technologies aim to extract value from data, they operate at different levels of sophistication and automation.

#### Business Intelligence Platforms: The Foundation

Business intelligence platforms serve as comprehensive systems that collect, integrate, and visualize data from [multiple sources](https://www.salesforce.com/uk/resources/articles/business-intelligence-platforms/?bc=HA). They excel at aggregating information from customer relationship management systems, enterprise resource planning software, sales databases, and other structured data sources. BI platforms transform raw data into dashboards, reports, and visualizations that help users understand what has happened in their business.

These platforms typically follow a structured process: gathering data from various sources, cleaning and organizing it, analyzing historical trends, and presenting findings through charts, graphs, and [dashboards](https://www.qlik.com/us/business-intelligence-platform). Users can explore data through pre-built reports or create custom visualizations to answer specific business questions.

#### Artificial Intelligence: The Evolution

AI represents a significant leap beyond traditional BI capabilities. While BI platforms require human analysts to interpret data and draw conclusions, AI systems can [automatically identify patterns, generate insights, and make predictions](https://www.oracle.com/be/artificial-intelligence/artificial-intelligence-analytics/). This fundamental difference transforms the relationship between humans and data from manual analysis to automated intelligence.

AI-powered analytics can process both structured and unstructured data, including text, images, and sensor data that traditional BI systems struggle to handle. More importantly, AI systems learn and improve over time, becoming more accurate and insightful as they process more data.

#### The Practical Differences

**Analysis Depth**: BI platforms excel at descriptive analytics—summarizing what happened in the past. AI systems provide predictive analytics (what will happen) and prescriptive analytics (what should be done).

**Automation Level**: BI platforms automate data collection and visualization but rely on human interpretation. AI systems can automate the entire analytical process, from data processing to insight generation and recommendation development.

**Data Handling**: BI platforms work primarily with structured data from known sources. AI systems can process diverse data types, including unstructured text, images, and [real-time sensor data](https://www.teradata.com/insights/ai-and-machine-learning/ai-for-data-analysis-and-visualization).

**User Interaction**: BI platforms require users to navigate dashboards and reports to find answers. AI systems can respond to [natural language queries](https://getdot.ai), making data access more intuitive and [democratic](https://dang.ai/tool/ai-data-analysis-getdot-ai).

#### The Integration Reality

Rather than replacing business intelligence platforms, AI enhances them. Modern BI platforms increasingly incorporate AI capabilities, creating hybrid systems that combine structured reporting with [intelligent analysis](https://www.oracle.com/be/artificial-intelligence/artificial-intelligence-analytics/). This integration provides the best of both worlds: the reliability and governance of traditional BI with the insight generation and automation of AI.

Organizations often maintain both systems, using BI platforms for routine reporting and governance while leveraging AI for complex analysis and prediction. This approach ensures data quality and compliance while enabling advanced analytics capabilities.

### AI in the Overall Data Analytics Ecosystem

Artificial intelligence occupies a crucial position within the broader data analytics ecosystem, serving as both an enhancement to existing capabilities and a transformative force for future development. Understanding this positioning helps organizations develop comprehensive data strategies that leverage AI's strengths while maintaining robust data governance and quality.

#### The Modern Data Analytics Stack

Today's data analytics ecosystem consists of multiple interconnected layers. At the foundation lies data infrastructure—data warehouses, lakes, and streaming platforms that store and manage organizational data. Above this sits the processing layer, including ETL tools, data preparation systems, and transformation engines that clean and organize data for analysis.

The analytics layer traditionally included business intelligence platforms, statistical analysis tools, and reporting systems. AI now augments this layer, adding capabilities for pattern recognition, predictive modeling, and automated insight generation. At the top, the presentation layer includes dashboards, reports, and increasingly, conversational interfaces that make [data accessible to all users](https://getdot.ai).

#### AI as an Enhancement Layer

Rather than replacing existing analytics infrastructure, AI functions as an enhancement layer that amplifies capabilities across the [entire stack](https://www.ibm.com/think/topics/ai-analytics). AI-powered data preparation tools can automatically clean and transform data, reducing the manual effort required for [analysis](https://www.teradata.com/insights/ai-and-machine-learning/ai-for-data-analysis-and-visualization). Machine learning algorithms can identify data quality issues, suggest corrections, and even predict data availability problems before they impact analysis.

In the analytics layer, AI transforms both the speed and depth of analysis possible. Where traditional analytics might take days or weeks to identify trends, AI systems can [process data in real-time, surfacing insights as they emerge](https://www.oracle.com/be/artificial-intelligence/artificial-intelligence-analytics/). This capability is particularly valuable for organizations dealing with high-velocity data streams from IoT devices, social media, or financial markets.

#### The Value Creation Process

AI's position in the analytics ecosystem creates value through several mechanisms. First, it democratizes access to sophisticated analysis by enabling [natural language queries and automated insight generation](https://getdot.ai). Second, it scales analytical capabilities beyond human limitations, processing vast amounts of data and identifying patterns that might be missed by traditional approaches.

<figure><img src="/files/cQXUkPuOeVxLdKmykgpJ" alt=""><figcaption></figcaption></figure>

*This diagram shows how AI integrates within the data analytics ecosystem, enhancing capabilities across multiple layers.*

#### Integration Challenges and Opportunities

Successful AI integration requires careful consideration of data quality, governance, and security. AI systems are only as good as the data they process, making robust [data governance essential](https://www2.deloitte.com/us/en/pages/consulting/articles/challenges-of-using-artificial-intelligence.html). Organizations must ensure that their data infrastructure can support AI workloads while maintaining security and [compliance requirements](https://docs.aws.amazon.com/prescriptive-guidance/latest/strategy-enterprise-ready-gen-ai-platform/best-practices.html).

The integration also presents opportunities for innovation. AI can help organizations discover new data sources, identify previously unknown relationships, and develop novel analytical approaches. Companies that successfully integrate AI into their analytics ecosystem often find new revenue streams and competitive advantages.

#### The Future of AI-Driven Analytics

The analytics ecosystem continues to evolve toward more intelligent, autonomous systems. Gartner predicts that by 2027, 75% of new analytics content will be contextualized through generative AI, enabling more dynamic and [automated decision-making](https://www.technologydecisions.com.au/content/security/news/the-near-future-of-analytics-in-the-ai-era-1603212566). This evolution suggests that AI will become increasingly central to the analytics ecosystem, ultimately becoming the primary interface between humans and [data](https://news.microsoft.com/source/features/ai/6-ai-trends-youll-see-more-of-in-2025/).

Organizations that understand AI's role in this broader ecosystem can make more strategic decisions about technology investments, skill development, and organizational capabilities. The key is recognizing that AI is not a standalone solution but part of a comprehensive approach to data-driven decision-making.

### Typical Use Cases and Applications

Artificial intelligence has transcended experimental applications to become a practical tool driving value across virtually every industry and [business function](https://research.aimultiple.com/ai-usecases/). Understanding these real-world applications helps organizations identify opportunities for AI implementation and measure potential returns on investment.

#### Customer Experience and Service

AI revolutionizes customer interactions through intelligent automation and personalization. Chatbots and virtual assistants now handle routine inquiries 24/7, resolving customer issues faster than traditional support [channels](https://builtin.com/artificial-intelligence/artificial-intelligence-in-business). Companies use AI to predict customer needs, recommending products before customers realize they want them.

Advanced customer service applications include sentiment analysis of support interactions, enabling companies to identify dissatisfied customers proactively and route complex issues to appropriate specialists. AI-powered systems can analyze customer communication patterns to predict churn risk, allowing businesses to intervene with targeted [retention strategies](https://www.upwork.com/resources/best-ai-use-cases).

#### Marketing and Sales Optimization

AI transforms marketing from broad campaigns to hyper-personalized experiences. Recommendation engines analyze customer behavior patterns to suggest products, content, or services with remarkable accuracy. Companies use AI to curate personalized content recommendations, dramatically improving user engagement and [satisfaction](https://www.upwork.com/resources/best-ai-use-cases).

Sales forecasting benefits enormously from AI's predictive capabilities. AI systems can analyze historical sales data, market conditions, and customer behavior to predict future demand with unprecedented accuracy. This capability helps organizations optimize inventory levels, plan production schedules, and allocate resources more effectively.

#### Financial Services and Risk Management

Financial institutions leverage AI for fraud detection, credit scoring, and algorithmic trading. AI systems can identify suspicious transaction patterns in real-time, flagging potentially fraudulent activities before they cause damage. Banks use machine learning to assess credit risk more accurately than traditional scoring methods, expanding access to credit while [reducing default rates](https://robertsmith.com/blog/applications-of-artificial-intelligence/).

Investment firms employ AI for portfolio optimization and market analysis. AI systems can process vast amounts of market data, news, and economic indicators to identify investment opportunities and manage risk more effectively.

#### Healthcare and Life Sciences

Healthcare represents one of AI's most promising application areas. AI systems assist with medical diagnosis by analyzing imaging data, identifying patterns that might be missed by human practitioners. Predictive models help healthcare providers anticipate patient needs, prevent complications, and optimize [treatment plans](https://redresscompliance.com/top-30-real-life-ai-use-cases-across-industries/).

Drug discovery benefits from AI's ability to analyze molecular structures and predict drug interactions. AI systems can identify potential drug candidates and predict their effectiveness, accelerating the development process and reducing costs.

#### Manufacturing and Operations

Manufacturing embraces AI for predictive maintenance, quality control, and process optimization. AI systems monitor equipment performance, predicting failures before they occur and scheduling maintenance to minimize [downtime](https://www.ibm.com/think/topics/artificial-intelligence-business-use-cases). Quality control systems use computer vision to identify defects faster and more accurately than human [inspectors](https://redresscompliance.com/top-30-real-life-ai-use-cases-across-industries/).

Supply chain optimization represents another significant AI application. AI systems analyze demand patterns, supplier performance, and logistics data to optimize inventory levels, reduce costs, and improve delivery times.

#### Human Resources and Talent Management

AI streamlines recruitment and talent management processes. AI-powered applicant tracking systems can screen resumes, identify qualified candidates, and even predict job performance based on historical data. However, organizations must carefully address bias concerns to ensure [fair and equitable hiring practices](https://www.ibm.com/think/insights/10-ai-dangers-and-risks-and-how-to-manage-them).

Employee engagement and retention benefit from AI's analytical capabilities. AI systems can analyze employee communication patterns, performance data, and other indicators to identify flight risks and suggest interventions.

#### Emerging Applications

New AI applications continue to emerge across industries. Smart cities use AI to optimize traffic flow, manage energy consumption, and improve public safety. Agriculture leverages AI for crop monitoring, pest detection, and yield [optimization](https://www.youtube.com/watch?v=KfRFClTRYVk). Even creative industries use AI for content generation, design optimization, and audience engagement.

The key to successful AI implementation lies in identifying specific business problems where AI's capabilities—pattern recognition, prediction, and automation—can provide measurable value. Organizations should start with well-defined use cases, establish clear success metrics, and gradually expand AI applications as they build expertise and confidence.

### What to Look Out for When Using or Buying AI Solutions

Implementing AI solutions requires careful consideration of multiple factors beyond basic functionality. Organizations that approach AI adoption strategically, with clear objectives and realistic expectations, achieve better outcomes than those rushing into implementation without proper planning.

#### Establishing Clear Business Objectives

The most common AI implementation pitfall involves investing in technology without a clear business case. Organizations should begin by identifying specific business problems rather than searching for AI applications. Successful AI initiatives start with questions like "How can we reduce customer churn?" or "What causes production delays?" rather than "How can we use AI?".

Clear objectives enable proper success measurement and ROI calculation. Organizations should define specific, measurable outcomes before implementation, such as reducing customer service response times by 40% or improving demand forecasting accuracy by 25%.

#### Data Quality and Governance

AI systems are only as good as the data they process, making data quality the foundation of successful [AI implementation](https://blog.dataiku.com/ai-gotchas-how-to-avoid-them). Poor data quality represents the primary reason AI projects fail. Organizations must audit their data sources, identify quality issues, and implement robust data governance practices before [deploying AI systems](https://www2.deloitte.com/us/en/pages/consulting/articles/challenges-of-using-artificial-intelligence.html).

Data governance becomes particularly critical when AI systems make decisions that impact customers or business operations. Organizations need clear policies for data collection, storage, and usage, along with audit trails that explain how AI systems [reach their conclusions](https://www.ibm.com/think/insights/ai-adoption-challenges).

#### Avoiding Bias and Ensuring Fairness

AI systems can inadvertently perpetuate or amplify existing biases present in training data. This risk is particularly significant in applications involving human decisions, such as hiring, lending, or law enforcement. Organizations must implement bias detection and mitigation strategies throughout the AI development lifecycle.

Regular testing and monitoring help identify bias in AI systems before they cause harm. Organizations should establish diverse teams to review AI systems and implement checks and balances to ensure fair outcomes across different [populations](https://mitsloanedtech.mit.edu/ai/basics/addressing-ai-hallucinations-and-bias/).

#### Infrastructure and Integration Requirements

AI workloads often require significant computational resources and specialized infrastructure. Organizations must assess their current technical capabilities and plan for necessary upgrades. Cloud-based AI services can reduce infrastructure requirements, but organizations must consider data security and [compliance implications](https://docs.aws.amazon.com/prescriptive-guidance/latest/strategy-enterprise-ready-gen-ai-platform/best-practices.html).

Integration with existing systems represents another critical consideration. AI solutions must work seamlessly with current data sources, business processes, and user workflows. Organizations should prioritize solutions that offer robust integration capabilities and consider the total cost of ownership, including integration and maintenance efforts.

#### Vendor Selection and Evaluation

The AI vendor landscape includes everything from established enterprise software companies to specialized AI startups. Organizations should evaluate vendors based on their track record, technical capabilities, and industry expertise. Proof-of-concept projects can help validate vendor claims and ensure solutions meet specific [business requirements](https://www.fastcompany.com/91183417/harnessing-ai-best-practices-for-genai-adoption-to-drive-company-growth).

Vendor stability and long-term viability represent important considerations, particularly for mission-critical applications. Organizations should assess vendors' financial stability, customer base, and product roadmap to ensure long-term support and [development](https://www.forbes.com/sites/delltechnologies/2025/05/27/from-good-to-growth-a-leaders-guide-to-boosting-ai-adoption/).

#### Security and Privacy Considerations

AI systems often process sensitive business and customer data, making security a paramount concern. Organizations must implement robust security measures to protect AI systems from attacks and ensure data privacy. This includes encryption, access controls, and regular [security audits](https://techinformed.com/5-strategies-for-responsible-ai-adoption/).

Regulatory compliance adds another layer of complexity. Organizations must ensure their AI systems comply with relevant regulations, such as GDPR for data privacy or [industry-specific requirements](https://www.forbes.com/sites/eliamdur/2023/09/13/pitfalls-in-artificial-intelligence-at-least-for-now/). Documentation and explainability features help demonstrate compliance and enable audit processes.

#### Change Management and User Adoption

Successful AI implementation requires significant organizational change. Users must understand new processes, learn new tools, and adapt to AI-enhanced workflows. Organizations should invest in comprehensive training programs and change management initiatives to ensure [successful adoption](https://www.forbes.com/sites/delltechnologies/2025/05/27/from-good-to-growth-a-leaders-guide-to-boosting-ai-adoption/).

Communication represents a critical success factor. Organizations should clearly explain AI capabilities and limitations to users, addressing concerns about job displacement and ensuring realistic expectations. Transparent communication about AI's role in decision-making helps [build trust and acceptance](https://techinformed.com/5-strategies-for-responsible-ai-adoption/).

#### Continuous Monitoring and Improvement

AI systems require ongoing monitoring and maintenance to ensure continued effectiveness. Model performance can degrade over time as data patterns change, requiring regular retraining and updates. Organizations should establish processes for monitoring AI system performance and implementing improvements.

Feedback mechanisms help identify issues and opportunities for enhancement. Organizations should create channels for users to report problems and suggest improvements, using this feedback to refine AI systems and processes.

### How AI-Driven Analytics Relates to the Future

The trajectory of AI-driven analytics points toward increasingly autonomous, intelligent systems that will fundamentally reshape how organizations interact with [data and make decisions](https://www.morganstanley.com/insights/articles/ai-trends-reasoning-frontier-models-2025-tmt). Understanding these emerging trends helps organizations prepare for the future and make strategic technology investments.

#### The Evolution Toward Autonomous Analytics

Current AI analytics systems primarily augment human decision-making by providing insights and recommendations. The future promises autonomous analytics platforms that can manage and execute business processes [independently](https://www.technologydecisions.com.au/content/security/news/the-near-future-of-analytics-in-the-ai-era-1603212566). Gartner predicts that by 2027, autonomous analytics will fully manage 20% of business processes, handling everything from data collection to [decision implementation](https://news.microsoft.com/source/features/ai/6-ai-trends-youll-see-more-of-in-2025/).

This evolution represents a shift from "analytics as a service" to "analytics as an agent". Future AI systems will proactively monitor business environments, identify opportunities and threats, and recommend or implement actions [automatically](https://www.technologydecisions.com.au/content/security/news/the-near-future-of-analytics-in-the-ai-era-1603212566). For example, AI agents might automatically adjust pricing strategies based on market conditions, optimize supply chains in response to demand changes, or reallocate marketing budgets to [maximize ROI](https://news.microsoft.com/source/features/ai/6-ai-trends-youll-see-more-of-in-2025/).

#### Multimodal AI and Enhanced Capabilities

Next-generation AI systems will seamlessly integrate and process multiple data types—text, images, audio, and video—in ways that mirror [human cognitive capabilities](https://www.informationweek.com/machine-learning-ai/what-will-be-the-next-big-thing-in-ai-). This multimodal approach will enable more sophisticated analysis and decision-making. Marketing teams might analyze customer sentiment from social media posts, video content, and voice interactions simultaneously to develop comprehensive customer insights.

Advanced reasoning capabilities will enable AI systems to solve complex problems through logical steps similar to human thinking. These systems will compare contracts, generate code, and execute multi-step workflows with minimal human oversight. The integration of these capabilities will make AI systems more versatile and valuable across [diverse business applications](https://www.informationweek.com/machine-learning-ai/what-will-be-the-next-big-thing-in-ai-).

#### The Rise of Conversational Analytics

The future of data interaction lies in natural language interfaces that make analytics accessible to everyone. Rather than navigating complex dashboards or writing queries, users will simply ask questions and receive comprehensive answers. This democratization of analytics will enable data-driven decision-making at every organizational level.

Tools like Dot, the AI data analyst, already demonstrate this future, enabling users to ask complex business questions in natural language and receive instant, accurate [insights](https://getdot.ai). As these interfaces become more sophisticated, they will handle increasingly complex analytical tasks, from root cause analysis to strategic [planning recommendations](https://dang.ai/tool/ai-data-analysis-getdot-ai).

#### AI-Powered Data Ecosystems

Future data ecosystems will be inherently intelligent, with AI embedded throughout the data lifecycle. AI will automatically discover new data sources, assess data quality, and optimize data flows for maximum value. These systems will continuously learn and adapt, becoming more effective over [time](https://news.microsoft.com/source/features/ai/6-ai-trends-youll-see-more-of-in-2025/).

Edge computing will enable real-time AI processing closer to data sources, reducing latency and enabling immediate decision-making. This capability will be particularly valuable for IoT applications, autonomous vehicles, and real-time fraud detection.

#### The Transformation of Business Models

AI-driven analytics will enable new business models based on data monetization and service [automation](https://www.morganstanley.com/insights/articles/ai-trends-reasoning-frontier-models-2025-tmt). Companies will offer AI-powered services that continuously adapt to customer needs, creating new revenue streams and competitive advantages. The subscription economy will expand to include AI-powered analytics services that provide ongoing value rather than one-time insights.

Organizations will increasingly compete on their ability to leverage AI for [strategic advantage](https://www.morganstanley.com/insights/articles/ai-trends-reasoning-frontier-models-2025-tmt). Those that successfully integrate AI into their operations will gain significant competitive advantages through improved efficiency, better decision-making, and enhanced customer experiences.

#### Preparing for the AI-Driven Future

Organizations can prepare for this future by building strong data foundations, developing AI literacy across their workforce, and experimenting with AI-driven analytics tools. The key is to start with practical applications that deliver immediate value while building capabilities for more advanced implementations.

Successful organizations will balance AI capabilities with human expertise, recognizing that the most effective approach combines artificial intelligence with human creativity and [judgment](https://techinformed.com/5-strategies-for-responsible-ai-adoption/). This collaboration will enable organizations to leverage AI's analytical power while maintaining the strategic thinking and ethical considerations that require human insight.

The future of AI-driven analytics promises unprecedented opportunities for organizations willing to embrace these technologies thoughtfully and [strategically](https://www.morganstanley.com/insights/articles/ai-trends-reasoning-frontier-models-2025-tmt). By understanding these trends and preparing accordingly, organizations can position themselves to thrive in an increasingly AI-driven business environment.

### Conclusion

Artificial intelligence represents a transformative force that is reshaping how organizations understand and leverage their data. From its origins in the 1950s with pioneers like Alan Turing and John McCarthy to today's sophisticated systems that can analyze vast datasets and provide real-time insights, AI has evolved from theoretical concept to practical business necessity.

The fundamental difference between AI and traditional business intelligence lies in AI's ability to learn, predict, and automate decision-making processes. While BI platforms excel at historical reporting and visualization, AI systems can process unstructured data, identify hidden patterns, and provide prescriptive recommendations that [transform how organizations operate](https://www.oracle.com/be/artificial-intelligence/artificial-intelligence-analytics/).

The business case for AI adoption is compelling. Organizations implementing AI-driven analytics report significant improvements in decision-making speed, operational efficiency, and competitive positioning. From predictive maintenance in manufacturing to personalized customer experiences in retail, AI applications deliver measurable value across industries.

However, successful AI implementation requires careful planning, quality data governance, and realistic expectations. Organizations must address challenges including data quality, bias mitigation, and change management while building the infrastructure and capabilities needed for [long-term success](https://www.ibm.com/think/insights/ai-adoption-challenges).

Looking forward, AI-driven analytics will become increasingly autonomous and accessible. The future promises conversational interfaces that democratize data analysis, multimodal systems that process diverse data types, and autonomous agents that can execute business [processes independently](https://www.technologydecisions.com.au/content/security/news/the-near-future-of-analytics-in-the-ai-era-1603212566). Tools like Dot, the AI data analyst, already demonstrate this future by enabling natural language [interactions with complex datasets](https://getdot.ai).

For organizations ready to embrace this transformation, the opportunity is substantial. AI is not just about technology adoption—it's about reimagining how businesses operate, compete, and create value in an increasingly data-driven world. The organizations that successfully integrate AI into their analytics ecosystems will gain significant competitive advantages through improved efficiency, better decision-making, and enhanced customer experiences.

The journey toward AI-driven analytics begins with understanding the technology's capabilities and limitations, identifying specific business problems where AI can provide value, and building the foundation for successful implementation. As AI continues to evolve, it will become an increasingly integral part of how organizations understand their world and shape their future.


# Business Intelligence

### What Exactly Is Business Intelligence?

Business intelligence, commonly known as **BI**, refers to the technologies, processes, and strategies organizations use to collect, analyze, and visualize data from internal and external sources. At its core, BI transforms raw information—from sales figures to website analytics—into digestible insights presented through dashboards, reports, and interactive visualizations. These insights help businesses understand what’s happening within their operations, marketplace, and customer behaviors, all with the [goal of improving day-to-day and strategic decision-making](https://www.altexsoft.com/blog/complete-guide-to-business-intelligence-and-analytics-strategy-steps-processes-and-tools/). Modern BI is far broader and more accessible: self-service BI tools allow non-technical users to generate their own reports and explore data—democratizing analytics across the organization. In short, business intelligence is the mechanism that puts comprehensive, up-to-date, and relevant data within reach of decision-makers at every level.

### Why Do Organizations Need Business Intelligence?

Organizations implement BI for a simple reason: in the age of data abundance, effective use of information is a differentiator. Decision-makers face a landscape filled with complexity and constant change. Without BI, large volumes of valuable data—about customer preferences, supply chain performance, marketing effectiveness, and more—remain siloed or underutilized, leading to missed opportunities, inefficiency, and reactive (rather than proactive) management.

Business intelligence provides several critical benefits:

**Informed Decision-Making and Reduced Guesswork**\
BI fosters an environment where choices are rooted in evidence, not intuition alone. Executives and managers can benchmark performance, spot risks early, and identify both strengths and inefficiencies quickly. For example, BI dashboards may [reveal geographic sales trends](https://www.altexsoft.com/blog/complete-guide-to-business-intelligence-and-analytics-strategy-steps-processes-and-tools/) enabling localized marketing campaigns, or highlight [slow-moving inventory](https://sproutsocial.com/insights/business-intelligence/) requiring proactive management.

**Comprehensive Business Visibility**\
By collecting, aggregating, and visualizing data from across departments—finance, sales, operations, marketing, customer service—BI unifies organizational intelligence. This holistic view allows leaders to align strategy, monitor key performance indicators (KPIs), and react to changes in the market or within internal operations in near real time.

**Competitive Advantage**\
A structured BI approach helps companies anticipate market shifts, benchmark against competitors, and discover new growth areas. For example, competitive BI may analyze public sentiment or rival pricing strategies, empowering an organization to seize emerging opportunities.

**Efficient Resource Allocation and ROI Optimization**\
BI allows firms to track which initiatives deliver returns and which expenses might be reallocated for greater impact. By tying spending to measurable outcomes, organizations can continually refine budgets and investments, ensuring resources flow to what genuinely works.

**Proactive Problem Solving and Innovation**\
When properly implemented, BI systems surface anomalies, risks, and trends before they escalate. Early-warning capabilities enable organizations to act faster, avoid costly mistakes, and foster a culture of continuous improvement. Insights from BI can also spotlight new product opportunities or process innovations—the seeds of competitive reinvention.

**Data Democratisation**\
Modern BI platforms empower not just IT or analytics teams but users throughout an organization to [access, explore, and act on data](https://sproutsocial.com/insights/business-intelligence/). Sales managers can analyze customer segments, HR can examine employee engagement stats, and operations teams can optimize workflows—all without waiting on central reporting teams.

**Strategic Alignment**\
A major benefit of BI is its role in aligning business units around shared, transparent metrics and objectives. By defining and tracking KPIs through BI dashboards, organizations reduce “silo” thinking and encourage all teams to row in the same direction.

*In essence, organizations need business intelligence because it is the backbone of a data-driven culture. BI sharpens awareness, shortens reaction times, and maximizes both the efficiency and impact of every decision.*

### BI Tools: Reporting, Dashboards, and Self-Service Analytics

BI technology comes in a few key flavors, each serving different needs:

* **Reporting tools** automate the [assembly of structured, regular updates](https://sproutsocial.com/insights/business-intelligence/)—think weekly sales reports or compliance checklists.
* **Dashboards** display metrics in real time, using visual elements like charts, gauges, and maps to make data both accessible and actionable.
* **Self-service analytics** takes BI even further, offering intuitive interfaces where business users can [explore data, ask ad hoc questions, and create their own custom visualizations](https://sproutsocial.com/insights/business-intelligence/) without needing technical training.

Major vendors (like Microsoft Power BI, Tableau, Qlik, Looker, SAP, and AWS QuickSight) have converged on platforms blending all three capabilities, supporting widespread BI adoption outside of IT and helping business units unlock rapid insight generation.

<figure><img src="/files/vqI5IryQRkxKJ18JUpKs" alt=""><figcaption></figcaption></figure>

### Business Intelligence vs. Data Science

Though business intelligence and data science both aim to unlock value from data, their methods, tooling, and focus diverge. **BI** is traditionally descriptive—it looks at historical and some current data to answer what happened. [Data science, by contrast, is predictive and occasionally prescriptive](https://corporatefinanceinstitute.com/resources/business-intelligence/business-intelligence-vs-data-science/), asking what will happen and why, often using [machine learning, statistical modeling](https://geeksforgeeks.org/data-science/difference-between-data-science-and-business-intelligence/), and more advanced techniques.

BI tools are generally designed for business users and focus on structured data. Data science works with a wider mix of data, including [semi- and unstructured text, images, and sensor data](https://geeksforgeeks.org/data-science/difference-between-data-science-and-business-intelligence/), often requiring specialized skills and coding expertise. In practice, BI and data science increasingly overlap, but each serves unique needs: BI for operational insight and control, data science for deeper exploration and longer-term forecasting.

### The Evolution of Business Intelligence

BI’s roots stretch back decades, originally encompassing IT-driven reporting and analytics based on transactional databases. Over time, advances in software, cloud computing, and data visualization transformed BI into an agile, user-driven capability. Today’s systems prioritize speed, accessibility, and visual storytelling, enabling even non-specialist users to engage with data, ask questions, and share findings rapidly.

### BI in the Modern Data Analytics Ecosystem

Business intelligence sits alongside transactional systems (where raw data is generated), data integration and storage platforms (such as data warehouses or lakes), and more advanced analytics (like data science or AI-powered analytics platforms). BI provides the essential bridge: translating raw, complex data into insights that drive business value and help organizations understand both the “what” and the “so what” behind their data-driven decisions.

### Typical BI Use Cases and Applications

In the real world, BI powers a wide spectrum of use cases:

* **Sales and revenue analysis:** Revealing product trends, customer segments, and opportunities for upsell or cross-sell
* **Marketing performance:** Tracking return on advertising spend, monitoring digital engagement, and optimizing campaign targeting
* **Supply chain management:** Spotting bottlenecks, optimizing stock levels, and ensuring timely fulfillment
* **Financial oversight:** Ensuring accurate budgeting, spend control, and real-time visibility into key financial drivers
* **Customer experience:** Analyzing support tickets, satisfaction surveys, and churn rates to improve service

All these applications share a common DNA: using data to anticipate needs, improve processes, and fuel growth, tailored to the specifics of each industry and organization.

### What to Look for When Adopting a BI Platform

Organizations evaluating BI tools should prioritize:

* **Robust data integration:** Can the tool connect to all data sources you need?
* **Scalability:** Does it handle current and future data volumes?
* **Self-service capabilities:** How easily can non-technical users generate insights?
* **Security:** Is data governance in place?
* **Vendor ecosystem support:** Are integrations and community resources available?

User-friendliness and adaptability are as important as technical horsepower—adoption fails if business teams cannot use the system directly.

### BI, AI, and the Road Ahead

Business intelligence is now converging with artificial intelligence and automated analytics. The future points towards insights delivered not only from what happened, but also why it happened, with automated recommendations for what to do next. [AI-driven tools](https://getdot.ai) are emerging as virtual data analysts, capable of sifting immense volumes of data, identifying anomalies, and even conducting natural language dialog with users for exploratory analysis.


# Data Warehouse

## What Exactly Is a Data Warehouse?

A data warehouse is a specialised data-management system that consolidates information from many source systems, stores it in a consistent, historical format, and [optimises it for analytical queries rather than day-to-day transactions](https://www.oracle.com/database/what-is-a-data-warehouse/). Bill Inmon, often called the “father of data warehousing”, distilled the idea into four attributes: subject-oriented, integrated, non-volatile and time-variant. Ralph Kimball later offered a complementary, bottom-up view, describing the warehouse as “a copy of transaction data specifically structured for query and analysis”.

In practical terms, a warehouse serves as a long-term “single source of truth” that supports business intelligence, regulatory reporting and now machine-learning workloads by storing [cleansed, conformed data that stretches years into the past](https://azure.microsoft.com/en-us/resources/cloud-computing-dictionary/what-is-a-data-warehouse). Unlike operational databases, which are optimised for inserts and updates, warehouses are engineered to scan billions of rows quickly, aggregate them on the fly and return consistent answers to complex questions.

#### Diagram: End-to-End Data Flow

<figure><img src="/files/D2Gu1cSpJYN3YvA0hNAJ" alt=""><figcaption></figcaption></figure>

*Description: data flows from diverse sources into a raw landing zone, is transformed into the central warehouse, and finally feeds downstream marts, dashboards and AI agents such as* [*Dot*](https://getdot.ai)*.*

## Why Do Organisations Need a Data Warehouse?

A warehouse answers the perennial executive question, “can we trust these numbers?” by enforcing common definitions and lineage across disparate datasets. Consolidation reduces the effort of reconciling sales ledgers with marketing funnels or supply-chain metrics, because the warehouse [standardises currencies, calendars and customer identifiers in one place](https://www.fivetran.com/learn/benefits-of-data-warehouse).

Performance is another driver. Analytical workloads that might take hours on operational systems can execute in seconds on [column-oriented warehouse storage](https://www.hava.io/blog/what-is-amazon-redshift), thanks to parallel processing, partition pruning and compressed data formats. Cloud-native platforms such as Snowflake and BigQuery add [automatic scaling](https://cloud.google.com/learn/what-is-a-data-warehouse), so weekend reporting bursts no longer require permanent hardware overprovisioning.

Historical depth also matters. Because warehouses store snapshots over time instead of overwriting yesterday’s state, analysts can [measure trends, seasonality and cohort behaviour](https://www.oracle.com/database/what-is-a-data-warehouse/) that operational databases simply lose. Time-travel features in modern systems even let users [query data “as of” a past moment](https://docs.snowflake.com/en/user-guide/intro-key-concepts), supporting audit and compliance work.

Governance is equally compelling. Centralised metadata, access controls and role-based security reduce the risk of ad-hoc data extracts circulating in spreadsheets. Data-quality monitors flag anomalies at ingestion, and catalogue tools attach business glossaries that demystify column names for non-technical staff.

Finally, a well-modelled warehouse sets the stage for advanced analytics. Training a predictive model demands clean, labelled data; warehouses provide exactly that foundation, which [AI services can consume directly](https://www.datasciencecentral.com/data-warehousing-reinvented-using-the-ai-advantage/) without repeated wrangling.

## How Do the Main Warehouse Architectures Differ?

Two classic philosophies dominate textbooks. Inmon’s “corporate information factory” builds a normalised, enterprise-wide repository first, then spins off data marts for departmental needs. Kimball’s dimensional approach starts with those marts but uses conformed dimensions so they can later knit together into a coherent whole.

Cloud vendors introduced a third family: decoupled storage and compute. Snowflake famously separates persistent object storage from ephemeral “virtual warehouses”, allowing independent scaling of each layer. Amazon Redshift takes a [cluster approach, adding concurrency scaling nodes on demand](https://en.wikipedia.org/wiki/Amazon_Redshift), while Google BigQuery is serverless—[users are billed per query rather than for fixed capacity](https://en.wikipedia.org/wiki/BigQuery).

Hybrid “lakehouse” and “logical” patterns have gained traction too. A lakehouse overlays [open-format table storage with warehouse-style metadata and ACID guarantees](https://cloud.google.com/architecture/big-data-analytics/data-warehouse), aiming to serve both data-science and BI users from one location. Logical data warehouses virtualise multiple physical stores behind a single semantic layer, federating queries without moving data.

#### Diagram: Three-Tier Snowflake Reference

<figure><img src="/files/aEAk7TNT9dobMPXk7eSW" alt=""><figcaption></figcaption></figure>

*Description: Snowflake’s architecture separates storage, compute and a cloud-services layer that handles security, metadata and optimisation, enabling elastic scaling and pay-as-you-go economics.*

## What’s the Difference Between a Data Warehouse and a Data Lake?

A lake stores raw, often unstructured files in their native format, [deferring schema definition until query time](https://www.talend.com/de/resources/data-lake-vs-data-warehouse/). This flexibility suits data scientists exploring clickstreams or sensor logs. A warehouse, by contrast, [imposes a schema before loading](https://www.coursera.org/articles/data-lake-vs-data-warehouse), ensuring that every table meets governance rules before any analyst runs a query.

Because lakes mix everything from images to JSON, they excel at experimentation but risk devolving into “data swamps” if curation lags. Warehouses trade flexibility for reliability: business analysts and finance teams can run month-end reports with [confidence that definitions are stable](https://azure.microsoft.com/en-us/resources/cloud-computing-dictionary/what-is-a-data-warehouse). Many organisations use both, [landing data in a lake and pushing cleansed subsets into the warehouse via ELT pipelines](https://www.reddit.com/r/dataengineering/comments/skrkoj/what_is_difference_between_data_warehouse_and/).

## How Did We Get Here? A Brief History

The concept emerged in the late 1980s at IBM as the “information warehouse” and took shape through Inmon’s 1992 book. Kimball’s *The Data Warehouse Toolkit* (1996) [democratised dimensional modelling](http://www.r-5.org/files/books/computers/databases/warehouses/Ralph_Kimball_Margy_Ross-The_Data_Warehouse_Toolkit-EN.pdf) for practitioners.

Early warehouses were on-premises, expensive and limited by hardware procurement cycles. The 2010s brought cloud disruptors: [Google BigQuery](https://cloud.google.com/learn/what-is-a-data-warehouse) (public release 2011), [Amazon Redshift](https://en.wikipedia.org/wiki/Amazon_Redshift) (2013) and [Snowflake](https://docs.snowflake.com/en/user-guide/intro-key-concepts) (general availability 2015), each abstracting infrastructure so teams could focus on data rather than servers.

The current decade adds AI and automation: [self-tuning workloads, natural-language interfaces and vector search](https://www.firebolt.io/blog/the-future-of-data-warehousing-in-the-age-of-ai-5-key-trends-from-firebolt-forward) to support generative models. Vendors now pitch “data clouds” or “unified analytics platforms” rather than standalone warehouses, reflecting the [convergence of storage, streaming and machine learning](https://www.databricks.com/resources/webinar/data-warehousing-era-ai).

## Where Does the Warehouse Fit in the Analytics Ecosystem?

Think of the warehouse as the hub of a modern data stack. Upstream, extract-load-transform (ELT) tools such as Airbyte and Fivetran move data from SaaS applications into cloud storage, while [dbt builds modular, version-controlled transformations](https://getdbt.com/) inside the warehouse. Downstream, visualisation layers issue SQL to the warehouse for dashboards, whereas orchestration engines [schedule data pipelines and provide lineage graphs](https://www.wherescape.com/blog/data-warehouse-automation-according-to-gartner/) for governance.

In AI workflows, notebooks and AutoML frameworks increasingly query the warehouse directly, eliminating duplicate feature stores. Agents such as [Dot](https://getdot.ai), the AI data analyst, connect to Snowflake, BigQuery or Redshift and translate natural-language questions into SQL, [returning charts and explanatory text](https://docs.getdot.ai) inside Slack or Teams.

## What Are Typical Use Cases?

Retailers correlate point-of-sale data with loyalty-card histories to [personalise promotions](https://www.databricks.com/discover/data-warehouse). Banks detect fraud by [analysing card transactions across years](https://www.scalefree.com/blog/data-warehouse/ai-in-data-warehousing-principles-and-applications/), flagging anomalous patterns in near real time. Manufacturers monitor IoT sensor feeds to [predict equipment failures](https://estuary.dev/blog/what-is-real-time-data-warehouse/), blending time-series data with maintenance logs in the warehouse. Healthcare providers merge claims, electronic health records and scheduling data to optimise resource utilisation. Governments aggregate tax, customs and social-service data to [spot compliance risks and model policy impacts](https://en.wikipedia.org/wiki/Bill_Inmon).

## What Should Buyers and Architects Look Out For?

Key evaluation criteria include performance at scale, cost transparency, concurrency limits, data-type support (structured, semi-structured, geospatial), security certifications, ecosystem integrations and vendor lock-in posture. Architectural choices—shared-nothing clusters versus serverless, lakehouse versus warehouse—should align with workload patterns, team skills and governance maturity.

Data quality and modelling discipline remain non-negotiable. A cloud subscription does not absolve teams from defining dimensional hierarchies, surrogate keys or slowly changing dimensions; [neglect here simply moves chaos to the cloud](https://www.wiley-vch.de/de/fachgebiete/computer-und-informatik/the-data-warehouse-toolkit-978-1-118-53080-1).

Change-data-capture pipelines, version-controlled transformations and comprehensive testing frameworks ensure that new source feeds do not break existing reports. [Observability platforms track freshness, volume and schema drift](https://www.wherescape.com/blog/data-warehouse-automation-according-to-gartner/), alerting teams before executives spot discrepancies.

## How Do Data Warehouses Relate to AI Analytics and Where Are They Headed?

AI is infiltrating every layer. Query optimisers now use machine-learning models to choose execution plans, reducing both latency and cost. [Vector databases and similarity search capabilities](https://www.databricks.com/resources/webinar/data-warehousing-era-ai) are being folded into mainstream warehouses so that large-language-model applications can retrieve context efficiently.

Natural-language interfaces, pioneered by tools like [Dot](https://getdot.ai), lower the skill barrier: business users type “why did monthly recurring revenue dip in June?” and an agent decomposes the question, generates SQL, visualises the output and returns recommended actions, all while respecting role-based permissions.


# Data Lake


# Data Visualization


# Large Language Model


# Machine Learning

## What Exactly Is Machine Learning?

Machine learning is a subset of artificial intelligence that enables computers to learn from data and make predictions or decisions without being explicitly programmed for every specific task ([learn from data and make predictions](https://www.sap.com/germany/products/artificial-intelligence/what-is-machine-learning.html)). At its core, it uses algorithms to identify patterns in data, learn from these patterns, and apply this knowledge to new, unseen information ([identify patterns in data](https://builtin.com/machine-learning/machine-learning-basics)). Rather than following pre-written instructions, machine learning systems improve their performance through experience—the more data they process, the better they become at their designated tasks ([improve their performance through experience](https://www.iso.org/artificial-intelligence/machine-learning)).

Tom Mitchell provided a formal definition that captures this essence: “A computer program is said to learn from experience E with respect to some class of tasks T and performance measure P, if its performance at tasks in T, as measured by P, improves with experience E” ([learn from experience](https://en.wikipedia.org/wiki/Machine_learning)). This learning process transforms raw data into actionable intelligence, enabling organizations to automate complex decision-making, predict future outcomes, and discover insights that would be impossible to identify through traditional programming approaches.

<figure><img src="/files/TIyvpDnnZEWcs5JZfqUa" alt=""><figcaption></figcaption></figure>

*This diagram illustrates how raw data and examples feed into algorithms that recognize patterns and learn, ultimately producing trained models capable of making predictions and generating insights.*

## Why Organizations Need Machine Learning

Organizations today face an unprecedented challenge: how to extract meaningful value from exponentially growing data volumes while maintaining competitive advantage in rapidly evolving markets. Machine learning addresses this by transforming data from a static resource into a dynamic, intelligence-generating asset that drives measurable business outcomes ([drives measurable business outcomes](https://www.forbes.com/councils/forbestechcouncil/2023/07/25/the-power-of-machine-learning-the-business-impact-on-real-time-data/)).

The business case centers on three value propositions: automation of complex processes, enhanced decision-making, and the ability to scale human expertise. Some organizations achieve up to 25% increases in profits through dynamic pricing models and automated optimization, such as when Amazon updates product prices every 10 minutes using machine learning—50 times more frequently than competitors ([dynamic pricing models](https://www.projectpro.io/article/machine-learning-use-cases/476)).

Predictive analytics and risk management are among the most compelling applications. Financial institutions use machine learning for fraud detection, achieving detection rates of over 93% while reducing false positives by more than 60% ([fraud detection systems](https://c3.ai/introduction-what-is-machine-learning/economic-or-business-value/)). Similarly, predictive maintenance systems can forecast equipment failures before they occur, reducing maintenance costs and preventing costly downtime.

The democratization of data insights through machine learning platforms enables non-technical users to query data in natural language, receive instant insights, and make data-driven decisions without relying on overburdened data teams ([natural language queries](https://www.getdot.ai)). This accessibility multiplies the value of organizational data by empowering more stakeholders to leverage insights for strategic decisions.

However, realizing these benefits isn’t without challenges. Many enterprise AI initiatives report average ROIs below 6%, often due to poor data quality, inadequate infrastructure, talent shortages, and lack of clear business alignment ([enterprise AI ROI](https://www.techmonitor.ai/ai-and-automation/enterprise-ai-adoption-accelerates-roi-elusive/)). Organizations that succeed focus on solving specific business problems and aligning initiatives with clear objectives.

<figure><img src="/files/jJ2T3AHLVwv75aY2i1xI" alt=""><figcaption></figcaption></figure>

*This diagram shows how machine learning addresses core challenges through specific solutions, ultimately delivering outcomes that drive organizational success.*

## Types of Machine Learning

Machine learning encompasses three primary paradigms, each designed for different problem types and data:

**Supervised Learning** uses labeled training data to make predictions about new examples, similar to learning with a teacher. Common applications include email spam detection and credit scoring ([make predictions about new examples](https://www.geeksforgeeks.org/machine-learning/supervised-vs-reinforcement-vs-unsupervised/)).

**Unsupervised Learning** discovers hidden patterns and structures within unlabeled data, exemplified by customer segmentation, where algorithms identify distinct groups without predefined labels ([discover hidden patterns](https://www.pecan.ai/blog/3-types-of-machine-learning/)).

**Reinforcement Learning** learns through trial and error by interacting with an environment and receiving rewards or penalties, much like how humans learn complex tasks. Autonomous vehicles and game-playing AI systems use this approach to optimize behavior ([learn through trial and error](https://www.digitalregenesys.com/blog/types-of-machine-learning-in-artificial-intelligence)).

*Different data types lead to approaches optimized for prediction, pattern discovery, or decision making.*

<figure><img src="/files/GiJRZkvAvJI99QsGZGGo" alt=""><figcaption></figcaption></figure>

## Machine Learning vs Artificial Intelligence

Artificial intelligence encompasses any machines performing tasks requiring human intelligence—reasoning, learning, and problem-solving—while machine learning specifically focuses on algorithms that improve through data experience ([algorithms that improve through data experience](https://ai.engineering.columbia.edu/ai-vs-machine-learning/)). Think of AI as the destination and machine learning as one of the primary vehicles to get there.

Many business “AI” solutions are actually machine learning systems that learn from data rather than following static rules. True machine learning systems improve over time, whereas rule-based systems remain static unless manually updated ([improve over time](https://www.coursera.org/articles/machine-learning-vs-ai)).

Deep learning is a branch of machine learning that uses multi-layer neural networks to process complex data. It excels at tasks like image recognition and natural language processing but requires substantial computational resources and data ([process complex data](https://aws.amazon.com/compare/the-difference-between-artificial-intelligence-and-machine-learning/)).

<figure><img src="/files/LAkrSBj4wUaDaXj9fVzO" alt=""><figcaption></figcaption></figure>

*Machine learning sits within the broader AI ecosystem, powering applications like computer vision, NLP, and recommendation systems.*

## A Brief History of Machine Learning

The evolution of machine learning spans decades, marked by key breakthroughs:

* **1940s–1950s**: Early neural network models by Pitts and McCulloch laid the groundwork for artificial neurons, and Arthur Samuel coined “machine learning” in 1959 to describe computers learning without explicit programming ([learn without explicit programming](https://www.lightsondata.com/the-history-of-machine-learning/)).
* **1950s–1960s**: Rosenblatt’s perceptron (1957) demonstrated pattern recognition capabilities, generating excitement about machine learning’s potential.
* **1970s–1980s (AI Winter)**: Minsky and Papert’s critique of perceptrons tempered enthusiasm, yet the backpropagation algorithm was rediscovered in the 1980s, crucial for training neural networks.
* **1990s–2000s (Renaissance)**: Support vector machines and random forests emerged, shifting focus from rule-based to data-driven approaches.
* **2010s–Present (Deep Learning Revolution)**: The release of ImageNet and the success of AlexNet in 2012 demonstrated deep neural networks’ power, leading to transformer architectures and large language models that underpin today’s AI applications ([power of deep neural networks](https://www.techtarget.com/whatis/feature/History-and-evolution-of-machine-learning-A-timeline)).

<figure><img src="/files/8pQR7aJKlJNz97aSROit" alt=""><figcaption></figcaption></figure>

*This timeline highlights foundational theories through today’s AI-powered products.*

## Machine Learning in the Data Analytics Ecosystem

Machine learning enhances the traditional analytics stack—data collection, storage, processing, and visualization—by automating anomaly detection, feature engineering, and real-time insights ([automating anomaly detection](https://www.ibm.com/think/topics/machine-learning-pipeline)). Modern platforms integrate AutoML to democratize model building, enabling business users to create predictive models without deep technical expertise.

Streaming analytics is critical for real-time fraud detection, recommendations, and monitoring. Platforms like Kafka process streaming data so organizations can respond as events occur rather than hours later ([process streaming data](https://aws.amazon.com/what-is/apache-kafka/)). Feature stores centralize the creation, storage, and serving of features, ensuring consistency between training and production environments.

AI-native analytics platforms, such as [getdot.ai](https://getdot.ai), allow users to query data with natural language, automatically generate visualizations, and receive intelligent recommendations about patterns. This breaks down barriers between data and insights, empowering domain experts to interact directly with analytics capabilities.

<figure><img src="/files/H02mvQTLt0KBAeOBedeG" alt=""><figcaption></figcaption></figure>

*Machine learning integrates batch and streaming data to produce predictions, dashboards, alerts, and API-driven responses.*

## Key Use Cases and Applications

Machine learning transforms industries:

* **Financial Services**: Real-time fraud detection achieves over 95% accuracy while reducing false positives, and algorithmic trading analyzes news and social media to predict market movements ([fraud detection accuracy](https://www.ibm.com/think/topics/machine-learning-use-cases)).
* **Healthcare**: Convolutional neural networks diagnose skin cancer with over 95% accuracy, and predictive models optimize treatment protocols and resource allocation ([diagnose skin cancer](https://www.tableau.com/learn/articles/machine-learning-examples)).
* **Retail & E-commerce**: Recommendation engines boost conversion rates and customer lifetime value, and dynamic pricing adjusts to demand in real time—Amazon’s approach yields 25% higher profits ([boost conversion rates](https://www.erlang-solutions.com/blog/machine-learning-for-business-benefits-examples/)).
* **Manufacturing**: Predictive maintenance forecasts equipment failures, reducing downtime, while computer vision inspects products for defects with high speed and precision ([predict equipment failures](https://research.aimultiple.com/ai-usecases/)).
* **Marketing & CX**: Sentiment analysis uncovers brand perception, churn models identify at-risk customers, and lead scoring pinpoints high-value prospects ([identify at-risk customers](https://www.datacamp.com/blog/top-machine-learning-use-cases-and-algorithms)).
* **Cybersecurity**: Anomaly detection flags unusual network behavior, and automated incident response systems contain threats without human intervention ([flag unusual network behavior](https://willdom.com/blog/how-do-machine-learning-and-artificial-intelligence-technologies-help-businesses/)).

<figure><img src="/files/b6JF4w7qaq379OovgkOD" alt=""><figcaption></figcaption></figure>

*This diagram links industries to applications and outcomes, illustrating machine learning’s breadth.*

## Considerations for Adopting Machine Learning Platforms

Selecting the right platform requires assessing:

* **Data Infrastructure**: Ensure seamless integration with warehouses like Snowflake or BigQuery and robust data governance to prevent failures due to poor data quality ([seamless integration](https://neptune.ai/blog/ml-platform-guide)).
* **Team Skills & Readiness**: Platforms with low-code interfaces help organizations struggling to find qualified talent, democratizing model building ([low-code interfaces](https://cake.ai/blog/machine-learning-platforms)).
* **Scalability & Performance**: Cloud-native solutions offer elasticity but raise sovereignty concerns, while on-premises options provide control at higher cost ([cloud elasticity](https://pecan.ai/blog/best-ai-platforms-guide/)).
* **Integration & Workflow**: Look for API support, version control, collaboration features, and automated deployment pipelines to streamline operations ([automated deployment](https://cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning)).
* **Governance & Compliance**: Role-based access, audit trails, and model explainability help meet regulations like GDPR and HIPAA ([model explainability](https://serengetitech.com/business/3-things-to-consider-before-implementing-machine-learning/)).

Proof of concept programs focusing on clear business problems can validate platform capabilities before major investments, while avoiding vendor lock-in ensures long-term flexibility.

## The Future of Machine Learning and AI Analytics

Emerging trends include:

* **No-Code Platforms**: Enabling business users to create models through natural language, with projections showing 70% of new applications built this way by 2025 ([70% by 2025](https://graphite-note.com/machine-learning-trends/)).
* **Real-Time & Edge Computing**: Deploying models at the edge for split-second decisions in autonomous vehicles and industrial automation ([split-second decisions](https://estuary.dev/blog/ai-trends/)).
* **Generative AI & LLMs**: Automating report writing, code generation, and synthetic data creation, enhancing traditional analytics with natural language interfaces ([automating report writing](https://cloud.google.com/transform/101-real-world-generative-ai-use-cases-from-industry-leaders)).
* **Automated Machine Learning (AutoML)**: Future platforms will handle feature engineering, data preprocessing, and deployment, reducing development time from months to hours ([handle feature engineering](https://aws.amazon.com/blogs/enterprise-strategy/unlocking-the-business-value-of-machine-learning-with-organizational-learning/)).
* **Federated Learning**: Training models on distributed data without sharing raw information, crucial for privacy-sensitive industries ([privacy-preserving training](https://www.apheris.com/resources/blog/how-to-choose-the-best-federated-learning-platform)).
* **Ethical AI & Responsible Development**: Integrated bias detection, fairness metrics, and explainability will become standard features as regulations tighten ([bias detection](https://it-dimension.com/blog/how-machine-learning-delivers-business-value-how-to-leverage-ml-for-growth-and-efficiency-mlai/)).

<figure><img src="/files/u5ND9vm2dPK3cNxGaVD2" alt=""><figcaption></figcaption></figure>

*This diagram shows how enabling technologies drive the evolution from expert-driven, manual operations to democratized, real-time, and automated machine learning.*

## Conclusion

Machine learning has evolved from academic curiosity to business imperative, transforming how organizations extract value from data and make decisions. Its power lies in augmenting human judgment with data-driven insights, enabling better, faster decisions and sustainable competitive advantages. Successful implementations start with clear business objectives, robust data governance, and platforms aligned with organizational capabilities. The democratization of AI through no-code interfaces like [getdot.ai](https://getdot.ai) is expanding access to advanced analytics, making machine learning an integral part of everyday business processes. As organizations balance technological innovation with business acumen, they will lead the next wave of data-driven innovation and competitive advantage.


# Semantic Layer

## What Exactly Is a Semantic Layer?

Business intelligence has spent three decades trying to turn column names into concepts people recognize. The [**semantic layer**](https://www.databricks.com/glossary/semantic-layer) is the piece of software—or sometimes just the design specification—that performs this translation. Sitting between raw storage and consumption tools, it maps tables and fields to business terms such as “customer,” “net revenue,” or “churn,” exposes those terms to every downstream application, and enforces the logic used to calculate them. In other words, it is the shared business vocabulary of an organization rendered in computable form. Modern implementations usually expose that vocabulary through [**SQL-generating APIs**](https://cube.dev/blog/semantic-layer-and-ai-the-future-of-data-querying-with-natural-language) or [**analytical services**](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-analyst/semantic-model-spec) so that notebooks, dashboards, and AI agents all speak the same language.

Because the layer hides schema complexity behind a consistent model, analysts can query data without memorizing joins, and developers can change physical models without breaking every dashboard. The result is faster insight and fewer reconciliation meetings.

***

## Why Do Organizations Need a Semantic Layer?

Warehouses and data lakes have solved scale, but not meaning. As Gartner and McKinsey repeatedly note, [**inconsistent metric definitions**](https://www.mckinsey.com/capabilities/quantumblack/our-insights/capturing-value-from-your-customer-data) remain a top cause of delayed decisions and mistrust in analytics. A semantic layer addresses that friction in four ways:

1. **Standardizes definitions.** When “active customer” is encoded once and referenced everywhere, Marketing, Finance, and Product teams stop debating whose number is “right,” and more time is spent on analysis instead of validation.
2. **Enables governed self-service.** Security and row-level policies travel with the business model, letting non-technical staff explore data safely while preserving lineage and compliance through a [**governed self-service**](https://www.ibm.com/think/topics/semantic-layer) framework.
3. **Cuts time-to-insight.** Metrics cached or pre-aggregated in a headless BI engine can [**reduce query latency**](https://www.atscale.com/blog/the-metric-store-and-its-role-in-the-modern-data-stack/) by orders of magnitude, making interactive exploration feasible on multi-terabyte tables.
4. **Grounds AI.** Large-language-model agents hallucinate when schemas are opaque; [**aligning prompts with a vetted business ontology triples answer accuracy**](https://technosapien.substack.com/p/how-semantic-data-layers-make-genai) in recent benchmarks. Vendors from Snowflake to dbt now embed semantic views specifically to feed AI copilots.

The net effect is a virtuous cycle: consistent metrics build trust; trust unlocks wider usage; usage justifies further investment in quality data.

***

## A Diagram of Placement in the Stack

<figure><img src="/files/xwWoV5lRJKHrykA6QUFo" alt=""><figcaption></figcaption></figure>

*This diagram shows raw data sources flowing into a central store; the semantic layer sits on top of that store and feeds every downstream interface, ensuring they all use the same definitions.*

***

## Different Types of Semantic Layer

Early layers lived inside proprietary BI platforms such as SAP BusinessObjects Universes. The modern “universal” layer is [**decoupled from any single tool**](https://cube.dev/blog/universal-semantic-layer-capabilities-integrations-and-enterprise-benefits) and exposed via open APIs. Some organizations embed the layer directly in the warehouse using features like [**Snowflake semantic views**](https://www.snowflake.com/en/engineering-blog/native-semantic-views-ai-bi/) while others choose external headless BI engines (Cube, AtScale, MetricFlow). For highly regulated workloads, a knowledge-graph-based semantic layer adds explicit ontologies and reasoning.

***

## Metric Store versus Semantic Layer

A metric store persists pre-computed numbers; a semantic layer encodes the logic to derive them on demand. In practice the metric store is often a service inside the broader semantic platform that [**materializes heavy aggregations**](https://www.kyvosinsights.com/blog/a-metrics-store-in-the-semantic-layer-architecture/) for performance, but it is not the whole story.

***

## A Short History

BusinessObjects patented the idea in 1991 as a way to shield users from SQL. Cognos, MicroStrategy, and others extended the concept in the 2000s. Cloud warehouses revived interest once data outgrew tightly coupled BI models; [**dbt Labs**](https://www.getdbt.com/blog/semantic-layer-introduction) brought metrics into version-controlled code, and [**Snowflake’s native semantic views**](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-analyst/semantic-model-spec) moved the layer even closer to the data. The present wave emphasizes tool-agnostic definitions and AI readiness.

***

## How the Semantic Layer Fits into the Data Analytics Ecosystem

In a modern stack, the layer complements transformation tools (dbt, Spark) by holding business logic that should not live in SQL pipelines. It feeds BI (Tableau, Power BI), operational analytics (reverse ETL), and LLMs via a single interface, acting as the contract between storage and every consumer. When combined with [**data catalogues**](https://www.collibra.com/us/en/blog/collibra-and-dbt-driving-a-common-language-around-data) from Collibra or Alation, it inherits lineage and governance metadata, closing the loop from source to insight.

***

## Typical Use Cases and Applications

* Executive dashboards that must reconcile revenue across sales channels.
* Self-service exploration for regional teams without exposing raw schemas.
* [**AI assistants**](https://getdot.ai) that [**translate a Slack question into validated SQL**](https://docs.getdot.ai) via the layer, removing hallucinations and surfacing explanations.
* Data-product APIs that expose governed metrics to partner ecosystems.
* Regulatory reporting where metric lineage and definitional control are audit requirements.

***

## What to Look Out for When Adopting or Buying

Evaluate breadth of tool integrations, modeling language expressiveness, performance on large joins, governance features, and version control. Ask vendors how they prevent metric drift, handle slowly changing dimensions, and expose programmatic APIs for CI/CD. Organizations that skip change management often find the centralized model becomes a bottleneck instead of a bridge.

***

## Semantic Layers, AI Analytics, and the Road Ahead

Large-language models excel at pattern detection but require context. A [**semantic layer supplies that context**](https://codd.ai/blog/semantic-layers-ai-bi) as structured metadata, enabling “chat with your data” without guesswork. Warehouse vendors are converging on an architecture where [**semantic views live natively**](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-analyst/semantic-model-spec) alongside tables, feeding RAG-based copilots that answer in business language while still surfacing SQL for audit. Expect future layers to learn relationships automatically, test metric validity continuously, and surface anomalies before users ask.


# KPI Trees

analytics frameworks for each business model


# Selling Assets

**Definition (short).** You earn money by selling a product once and transferring ownership. After that upfront payment, obligations are limited (warranty, support). Classic manufacturing/retail: revenue = units × price.

**Recent examples.** Apple Inc. exemplifies asset sales: [**51% of Apple’s \~$391 billion FY 2024 revenue came from iPhone hardware**](https://www.businessofapps.com/data/apple-statistics/?utm_source=chatgpt.com), a one-time purchase product line. Automakers like Toyota and consumer electronics giants such as Samsung also rely on outright product sales as their core revenue engines.

**Historical example.** The Ford Model T sold [**over 15 million units between 1908 and 1927**](https://en.wikipedia.org/wiki/Ford_Model_T), proving that scale manufacturing plus a single upfront price could transform affordability—and a company’s economics.

<figure><img src="/files/2jAY0DD0hykjREbqxbET" alt=""><figcaption></figcaption></figure>

#### KPI Definitions (matches the nodes above)

1. **Profitable Growth (composite)**

   *EN:* Balanced growth in revenue with healthy profitability.

   *Pseudo:* `w1 * Revenue_Growth% + w2 * Net_Profit_Margin%` or `Grow revenue while NPM ≥ threshold`.

   *Why:* Forces trade-off clarity—no growth-at-all-costs or margin-at-all-costs blind spots.

   *Benchmark:* Exec teams often set explicit weights or guardrails (e.g., “≥10% growth AND NPM ≥20%”).
2. **Revenue Growth %**\
   \&#xNAN;*EN:* YoY % change in product revenue.\
   \&#xNAN;*Pseudo:* `(Rev_t − Rev_{t−1}) / Rev_{t−1} * 100`\
   \&#xNAN;*Why:* First read on demand, pricing power, and market share shifts. Sustained high growth buys strategic optionality.\
   \&#xNAN;*Benchmark:* Mature manufacturers often grow \~5–8% YoY, while top-decile durables can exceed 20% in expansion phases.
3. **Gross Margin %**\
   \&#xNAN;*EN:* Share of revenue retained after COGS.\
   \&#xNAN;*Pseudo:* `(Revenue − COGS) / Revenue * 100`\
   \&#xNAN;*Why:* Core unit economics; funds SG\&A, R\&D, profit.\
   \&#xNAN;*Benchmark:* Consumer electronics median ≈ 30%; Apple’s overall gross margin hit **\~46.9% in Q1 FY25**.
4. **Units Sold**\
   \&#xNAN;*EN:* Total items delivered in period.\
   \&#xNAN;*Pseudo:* `Σ units_sold`\
   \&#xNAN;*Why:* Volume driver; reveals penetration and lifecycle stage.\
   \&#xNAN;*Benchmark:* Flagship auto models sell \~1 M units/year; niche B2B devices sell in the thousands.
5. **Average Selling Price (ASP)**\
   \&#xNAN;*EN:* Realized average price per unit after discounts.\
   \&#xNAN;*Pseudo:* `Revenue / Units_Sold`\
   \&#xNAN;*Why:* Signals pricing power and mix (premium vs entry). Rising ASP can offset flat volume.\
   \&#xNAN;*Benchmark:* iPhone ASP ≈ $800, while global smartphone ASP sits near $285–300.
6. **Net Profit Margin %**\
   \&#xNAN;*EN:* Net income as % of revenue.\
   \&#xNAN;*Pseudo:* `Net_Income / Revenue * 100`\
   \&#xNAN;*Why:* Bottom-line health; compresses strategy + execution into one number.\
   \&#xNAN;*Benchmark:* Premium hardware firms \~15–25%; big-box retail often 2–4%. Apple FY 24 \~24–25%.
7. **Operating Expense Ratio %**\
   \&#xNAN;*EN:* SG\&A + R\&D as % of revenue.\
   \&#xNAN;*Pseudo:* `Opex / Revenue * 100`\
   \&#xNAN;*Why:* Cost leverage; shows whether scale translates to profit.\
   \&#xNAN;*Benchmark:* Lean manufacturers 10–15%; many tech-heavy hardware players hover \~20%.
8. **Inventory Turnover (x)**\
   \&#xNAN;*EN:* Times inventory turns per year.\
   \&#xNAN;*Pseudo:* `COGS / Avg_Inventory`\
   \&#xNAN;*Why:* Working-capital efficiency; slow turns trap cash and risk obsolescence.\
   \&#xNAN;*Benchmark:* Durable goods 5–8x; fast fashion >10x.
9. **Number of Customers**\
   \&#xNAN;*EN:* Distinct purchasing customers in period.\
   \&#xNAN;*Pseudo:* `COUNT(DISTINCT customer_id)`\
   \&#xNAN;*Why:* Market breadth and concentration risk indicator.\
   \&#xNAN;*Benchmark:* B2C brands = millions; capital equipment vendors = dozens/hundreds.
10. **Units per Customer**\
    \&#xNAN;*EN:* Average quantity each customer buys.\
    \&#xNAN;*Pseudo:* `Units_Sold / Customers`\
    \&#xNAN;*Why:* Captures repeat purchase and bundle depth.\
    \&#xNAN;*Benchmark:* Big-ticket durables ≈1; consumables/accessories 3–5+.
11. **List Price**\
    \&#xNAN;*EN:* MSRP before discounts.\
    \&#xNAN;*Pseudo:* `MSRP_value`\
    \&#xNAN;*Why:* Psychological anchor; defines promo headroom.\
    \&#xNAN;*Benchmark:* Premium brands often list 20–50% above category medians.
12. **Discount Rate %**\
    \&#xNAN;*EN:* Average markdown from list to realized price.\
    \&#xNAN;*Pseudo:* `(List − Realized) / List * 100`\
    \&#xNAN;*Why:* Tracks promo dependency and margin leakage.\
    \&#xNAN;*Benchmark:* Consumer electronics promos 10–15%; luxury <5% except controlled clearance.
13. **COGS per Unit**\
    \&#xNAN;*EN:* Direct cost per unit.\
    \&#xNAN;*Pseudo:* `COGS / Units_Sold`\
    \&#xNAN;*Why:* Every dollar saved drops to gross margin; core lean initiative.\
    \&#xNAN;*Benchmark:* Targets of 2–5% YoY reduction are common in best-in-class ops.
14. **Days of Inventory (DOI)**\
    \&#xNAN;*EN:* Average days inventory sits before sale.\
    \&#xNAN;*Pseudo:* `365 / Inventory_Turnover`\
    \&#xNAN;*Why:* Cash velocity and obsolescence risk measure.\
    \&#xNAN;*Benchmark:* Best-in-class electronics <45 days; typical manufacturers 60–90 days.


# Leasing Assets

**Definition (short).** Customers pay for time-bound use of an asset (room, car, excavator, desk). You retain ownership; revenue hinges on utilization (how often it’s rented) and rate (what you charge per unit-time).

**Recent examples.** U.S. hotels posted [**record ADR ($155.62), record RevPAR ($97.97), and 63% occupancy in 2023**](https://www.traveldailynews.com/statistics-trends/u-s-hotels-posted-record-high-adr-and-revpar-in-2023/?utm_source=chatgpt.com), marking a full recovery in pricing and near-recovery in utilization. The U.S. car rental industry hit $38.3 billion revenue in 2023, up 5% YoY—another record.

**Historical example.** [Blockbuster](https://en.wikipedia.org/wiki/Blockbuster_LLC) rented movies/games for a few days: at its 2004 peak it ran \~9,000 stores and earned $5.9 billion—pure temporary access to physical media.

<figure><img src="/files/C52vUAoeC7HJflqN8N6x" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

1. **Profitable Utilization**

   *EN*: How much profit you generate for every unit of capacity you could rent (room-night, car-day, excavator-hour).

   *Pseudo*:

   * Profitable Utilization = RevPU \* EBIT\_Margin
   * or equivalently ((Revenue − Opex − Maintenance − Depreciation) / (Units \* Time))

   *Why*: Blends the two levers that matter most in renting: rate × fill (RevPU) and cost discipline (margin). One number forces trade-off clarity between pushing price/occupancy and controlling asset & operating costs.

   *Benchmark idea*: Hotels often track GOPPAR (gross operating profit per available room). This version just generalizes that: Profit per Available Unit-Time (PPAU). Set a target (e.g., grow PPAU ≥ X% YoY while occupancy stays between 70–85%).
2. **Revenue Growth %**\
   \&#xNAN;*EN:* YoY % change in lease/rental revenue.\
   \&#xNAN;*Pseudo:* `(Rev_t − Rev_{t−1}) / Rev_{t−1} * 100`\
   \&#xNAN;*Why:* Shows demand/pricing momentum, especially critical post-shocks (pandemics, recessions).\
   \&#xNAN;*Benchmark:* U.S. hotel RevPAR grew +4.9% in 2023; car rental revenue +5% YoY.
3. **Revenue per Available Unit (RevPAR / RPU)**\
   \&#xNAN;*EN:* Revenue divided by total unit-time available.\
   \&#xNAN;*Pseudo:* `Revenue / (Units * Time)` or `Occupancy * ADR`.\
   \&#xNAN;*Why:* Single metric blending rate and utilization; gold standard for asset productivity.\
   \&#xNAN;*Benchmark:* U.S. hotels averaged RevPAR $97.97 in 2023.
4. **Occupancy / Utilization %**\
   \&#xNAN;*EN:* % of available unit-time actually rented.\
   \&#xNAN;*Pseudo:* `Occupied_Time / Available_Time * 100`\
   \&#xNAN;*Why:* Idle assets = lost revenue; too high leaves no room for peak pricing.\
   \&#xNAN;*Benchmark:* U.S. hotel occupancy 63% in 2023; top-city properties routinely hit 80%+ in peak seasons.
5. **Average Rental Rate (ADR/ARR)**\
   \&#xNAN;*EN:* Average realized price per rented unit-time.\
   \&#xNAN;*Pseudo:* `Revenue / Rented_Time`\
   \&#xNAN;*Why:* Pricing power lever; raising ADR lifts RevPU without more assets.\
   \&#xNAN;*Benchmark:* U.S. hotel ADR $155.62 (2023); upscale segments far higher.
6. **EBIT Margin %**\
   \&#xNAN;*EN:* Operating profit (pre-interest/taxes) as % of revenue.\
   \&#xNAN;*Pseudo:* `EBIT / Revenue * 100`\
   \&#xNAN;*Why:* Captures structural profitability after maintenance, overhead.\
   \&#xNAN;*Benchmark:* Car rental majors often post 10–15% EBIT; asset-light hotel franchisors can exceed 30%.
7. **Maintenance & Depreciation Ratio %**\
   \&#xNAN;*EN:* Maintenance plus depreciation cost share of revenue.\
   \&#xNAN;*Pseudo:* `(Maintenance + Depreciation) / Revenue * 100`\
   \&#xNAN;*Why:* Assets wear out; keeps margin honest. Too low risks quality, too high erodes profits.\
   \&#xNAN;*Benchmark:* Fleet-heavy rentals \~9–12%; hotels allocate \~5–8% of revenue to property maintenance/capex.
8. **Contract / Lease Renewal %**\
   \&#xNAN;*EN:* Portion of expiring long-term leases/contracts that renew.\
   \&#xNAN;*Pseudo:* `Renewed / Expiring * 100`\
   \&#xNAN;*Why:* Stickiness metric; lowers selling cost and smooths cash flows.\
   \&#xNAN;*Benchmark:* Commercial leases 75–90% renewal; equipment leases 60–80%+ depending on term.
9. **Total Capacity (Units \* Time)**\
   \&#xNAN;*EN:* Theoretical rentable supply (rooms × nights, cars × days).\
   \&#xNAN;*Pseudo:* `Units * Available_Time`\
   \&#xNAN;*Why:* Sets denominator for utilization/RevPU; critical for planning expansion or retirements.
10. **Occupied / Rented Units**\
    \&#xNAN;*EN:* Actual unit-time rented in period.\
    \&#xNAN;*Pseudo:* `Σ rented_unit_time`\
    \&#xNAN;*Why:* Raw volume driver for utilization and revenue; track by segment to see mix shifts.
11. **Avg Rental Duration**\
    \&#xNAN;*EN:* Mean length of each rental/booking.\
    \&#xNAN;*Pseudo:* `Total_Rented_Time / #Contracts`\
    \&#xNAN;*Why:* Impacts ops (turn costs), pricing, and forecasting.\
    \&#xNAN;*Benchmark:* Cars 3–7 days; equipment months; hotels 1–3 nights average.
12. **Yield Management Uplift %**\
    \&#xNAN;*EN:* Incremental revenue vs flat pricing.\
    \&#xNAN;*Pseudo:* `(Actual_Rev − FlatRate_Rev) / FlatRate_Rev * 100`\
    \&#xNAN;*Why:* Quantifies value of RM/dynamic pricing systems.\
    \&#xNAN;*Benchmark:* Airlines/hotels typically claim +3–7% uplift from RM algorithms.
13. **Dynamic Pricing Hit Rate %**\
    \&#xNAN;*EN:* % of price changes that improved RevPU.\
    \&#xNAN;*Pseudo:* `#Positive_Changes / Total_Changes * 100`\
    \&#xNAN;*Why:* Ensures pricing AI is actually adding value, not noise.\
    \&#xNAN;*Benchmark:* Internal metric—teams target quarter-over-quarter improvement.


# Offering Subscription

**Definition (short).** Customers pay a recurring fee (monthly/annual) for continuous access to a product, service, or perks. Value comes from retention, expansion, and efficient acquisition—far more about relationships than one-off deals.

**Recent examples.** Netflix closed 2024 with [\~301.6 million paying members](https://apnews.com/article/c0447b9289e31e09ce4f0b6e6bde1c54?utm_source=chatgpt.com) and keeps raising ARPU via pricing tiers and an ad-supported plan. Spotify ended Q4 2024 at [263 million premium subscribers](https://newsroom.spotify.com/2025-02-04/spotify-reports-fourth-quarter-2024-earnings/?utm_source=chatgpt.com) and posted its first full profitable year. B2B SaaS leaders (Salesforce, ServiceNow) live and die by ARR growth and net retention.

**Historical example.** Newspaper and magazine subscriptions date to the 18th century; *The Times* (1785) and *Reader’s Digest* (1922) built massive recurring bases long before software did.

<figure><img src="/files/F0kVPQkFjWx1glqkl0GV" alt=""><figcaption></figcaption></figure>

1. **Net Recurring Profit Growth % (NRG)**\
   \&#xNAN;*EN:* Year-over-year growth in profit generated by recurring revenue (ARR × Gross Margin − OpEx tied to subs).\
   \&#xNAN;*Pseudo:*

   ```
   ((ARR_t*GM_t - OpEx_t) - (ARR_{t-1}*GM_{t-1} - OpEx_{t-1})) / (ARR_{t-1}*GM_{t-1} - OpEx_{t-1}) * 100  
   ```

   *Why:* Bundles what a CEO ultimately wants from a subscription engine: bigger recurring top line *and* healthy margins.\
   \&#xNAN;*Benchmark:* Private SaaS medians target **\~20% ARR growth** with **\~75–85% gross margin** ([OpenView](https://openviewpartners.com/2023-saas-benchmarks-report/?utm_source=chatgpt.com)).
2. **ARR / MRR Growth %**\
   \&#xNAN;*EN:* % change in annual or monthly recurring revenue.\
   \&#xNAN;*Pseudo:* `(ARR_t - ARR_{t-1}) / ARR_{t-1} * 100`\
   \&#xNAN;*Why:* Clean topline indicator, independent of one-off services.\
   \&#xNAN;*Benchmark:* 2024 SaaS ARR growth medians **\~19%**.
3. **Net Profit Margin %**\
   \&#xNAN;*EN:* Net income / revenue.\
   \&#xNAN;*Pseudo:* `NI / Revenue * 100`\
   \&#xNAN;*Why:* Recurring revenue still needs to drop profit to the bottom line.\
   \&#xNAN;*Benchmark:* Large mature subs businesses (e.g., Spotify) moved to positive NP margins; best-in-class SaaS aim double-digit NP margins at scale.
4. **Active Subscribers (Subs)**\
   \&#xNAN;*EN:* Paying customers at period end.\
   \&#xNAN;*Pseudo:* `Subs_end = Subs_start + New - Churned`\
   \&#xNAN;*Why:* The “count” lever of ARR (ARR = Subs × ARPU).\
   \&#xNAN;*Benchmark:* Netflix 301.6 M; Spotify 263 M; B2B SaaS counts vary—growth rate matters more than absolute.
5. **ARPU / ARPA**\
   \&#xNAN;*EN:* Average revenue per user/account.\
   \&#xNAN;*Pseudo:* `Recurring_Revenue / Avg_Subs`\
   \&#xNAN;*Why:* Monetization depth; can rise via price hikes, tiering, upsells.\
   \&#xNAN;*Benchmark:* Streaming ARPU spans **$4–$14/mo** globally; enterprise SaaS ARPA often **>$50k/yr**.
6. **Net Revenue Retention % (NRR)**\
   \&#xNAN;*EN:* (Starting ARR from existing customers ± expansion − contraction − churn) / Starting ARR.\
   \&#xNAN;*Pseudo:* `(ARR_existing_end / ARR_existing_start) * 100`\
   \&#xNAN;*Why:* Land-and-expand scorecard. >100% means base grows without new logos.\
   \&#xNAN;*Benchmark:* Public/private SaaS medians **\~101% NRR**, top quartile **>120%**.
7. **New Subscribers**\
   \&#xNAN;*EN:* Gross adds in period.\
   \&#xNAN;*Pseudo:* `count(new_customer_ids)`\
   \&#xNAN;*Why:* Pipeline health; offsets churn to grow the base.\
   \&#xNAN;*Benchmark:* Aim that New Subs ≥ Churned Subs to lift net adds.
8. **Churn % (Logo or Revenue)**\
   \&#xNAN;*EN:* % of customers (or ARR) lost in period.\
   \&#xNAN;*Pseudo:* `Churned_Customers / Start_Customers * 100` (logo) or `Lost_ARR / Start_ARR * 100`\
   \&#xNAN;*Why:* Leaky-bucket killer; drives LTV.\
   \&#xNAN;*Benchmark:* B2B SaaS best-in-class **<10% annual logo churn**; consumer subs **3–5% monthly** common.
9. **Expansion / Upsell Rev %**\
   \&#xNAN;*EN:* Additional ARR from existing customers / starting ARR.\
   \&#xNAN;*Pseudo:* `Expansion_ARR / Start_ARR * 100`\
   \&#xNAN;*Why:* Determines NRR >100%; cheaper than net-new.\
   \&#xNAN;*Benchmark:* Top PLG SaaS drive **15–30% expansion annually**.
10. **Base Plan Price**\
    \&#xNAN;*EN:* Published price for entry tier.\
    \&#xNAN;*Pseudo:* `Plan_price_monthly`\
    \&#xNAN;*Why:* Anchor for ARPU and upgrade ladders; careful adjustments yield ARPU gains without spiking churn.\
    \&#xNAN;*Benchmark:* Streaming base plans **$7–15**; SMB SaaS **$20–50**; enterprise tiers negotiated.
11. **Gross Retention % (GRR)**\
    \&#xNAN;*EN:* % of starting ARR retained *excluding* expansion.\
    \&#xNAN;*Pseudo:* `(ARR_existing_end_excl_expansion / ARR_existing_start) * 100`\
    \&#xNAN;*Why:* Purist retention health; insulation from upsell masking churn.\
    \&#xNAN;*Benchmark:* SaaS GRR median **\~90%**.
12. **Cross-sell Rate**\
    \&#xNAN;*EN:* % of customers buying ≥2 products/modules.\
    \&#xNAN;*Pseudo:* `Customers_multi_product / Total_Customers * 100`\
    \&#xNAN;*Why:* Depth of relationship; diversifies revenue per account.\
    \&#xNAN;*Benchmark:* Top enterprise suites push **>30% cross-sell penetration**.
13. **Customer Acquisition Cost (CAC)**\
    \&#xNAN;*EN:* Sales & marketing spend per new paying customer.\
    \&#xNAN;*Pseudo:* `S&M_spend_period / New_Paying_Customers`\
    \&#xNAN;*Why:* Efficiency of growth engine; ties to payback and LTV economics.\
    \&#xNAN;*Benchmark:* SaaS CAC often **$200–$600** for SMB, thousands for enterprise.
14. **Payback Months**\
    \&#xNAN;*EN:* Months to recover CAC from gross margin dollars.\
    \&#xNAN;*Pseudo:* `CAC / (ARPU * GM%)`\
    \&#xNAN;*Why:* Shorter payback = faster cash recycle, less burn.\
    \&#xNAN;*Benchmark:* Best-in-class SaaS **≤12 months**; median **15–16 months**.
15. **Customer Lifetime Value (LTV)**\
    \&#xNAN;*EN:* Gross margin dollars expected from a customer over life.\
    \&#xNAN;*Pseudo:* `(ARPU * GM%) / Churn_rate`\
    \&#xNAN;*Why:* Sets the ceiling for rational CAC.\
    \&#xNAN;*Benchmark:* Wildly variable; LTV/CAC ratio is key.
16. **LTV/CAC Ratio**\
    \&#xNAN;*EN:* LTV divided by CAC.\
    \&#xNAN;*Pseudo:* `LTV / CAC`\
    \&#xNAN;*Why:* Sanity check: too low = unprofitable growth; too high = under-investment.\
    \&#xNAN;*Benchmark:* “Rule of thumb” **3:1 ideal**, <2:1 dangerous, >5:1 may mean under-investing.


# Offering metered Usage

**Definition (short).** Customers pay in proportion to what they consume—API calls, GB stored, kWh, messages, transactions. Revenue scales with usage, not (only) with time or seats; pricing can be linear or tiered.

**Recent examples.** AWS, Azure, and GCP bill per compute hour, GB, or request; AWS alone produced >$100 B in 2024 revenue. Utilities charge [\~16.4¢ per kWh](https://www.eia.gov/electricity/monthly/epm_table_grapher.php?t=epmt_5_6_a\&utm_source=chatgpt.com) to U.S. residential customers on average (Apr 2025), and the average home uses [\~10,500 kWh/year](https://www.eia.gov/energyexplained/use-of-energy/electricity-use-in-homes.php?utm_source=chatgpt.com) (≈875 kWh/mo).

**Historical example.** Edison’s Pearl Street Station (1882) billed electricity by the kilowatt‐hour—one of the earliest metered models. Taxi meters (1890s) and postage (per-letter) are other classic pay-per-use forebears.

<figure><img src="/files/YRrfgUplQ86kybINLBZ6" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

1. **Profit per Unit of Usage (PPUU)**\
   \&#xNAN;*EN:* Profit generated for each billable unit (after variable costs).\
   \&#xNAN;*Pseudo:* `PPUU = (Effective_Unit_Price - Unit_Cost) / Unit` or more commonly `(Revenue - Variable_Costs) / Total_Units`

   *Why:* Forces clarity on both monetization (price) and efficiency (cost). High volume without profit per unit is a trap; high margin without volume underutilizes the model.\
   \&#xNAN;*Benchmark:* Digital infra aims **70–85% gross margin**.
2. **Total Usage Volume (UV)**\
   \&#xNAN;*EN:* Sum of consumed units (API calls, kWh, GB, etc.).\
   \&#xNAN;*Pseudo:* `Σ usage_units_all_customers`\
   \&#xNAN;*Why:* Primary revenue driver; shows engagement/demand intensity.\
   \&#xNAN;*Benchmark:* Target **>30–50% YoY** usage growth early on.
3. **Effective Unit Price (EUP)**\
   \&#xNAN;*EN:* Average realized price per unit after tiers/discounts.\
   \&#xNAN;*Pseudo:* `Total_Usage_Revenue / Total_Usage_Volume`\
   \&#xNAN;*Why:* Monetization efficiency; erosion may signal commoditization.
4. **Gross Margin per Unit % (GMU)**\
   \&#xNAN;*EN:* % of unit price kept after variable cost.\
   \&#xNAN;*Pseudo:* `(EUP - Unit_Cost) / EUP * 100`\
   \&#xNAN;*Why:* Unit economics; determines scalability of volume growth.
5. **Active Users / Accounts (AU)**\
   \&#xNAN;*EN:* Unique customers generating billable usage in period.\
   \&#xNAN;*Pseudo:* `COUNT(DISTINCT user_id WHERE usage>0)`\
   \&#xNAN;*Why:* Breadth of adoption; needed to diversify revenue and de-risk whale dependence.
6. **Average Usage per User (AUPU)**\
   \&#xNAN;*EN:* Mean consumption per active user.\
   \&#xNAN;*Pseudo:* `Total_Usage_Volume / Active_Users`\
   \&#xNAN;*Why:* Depth of engagement; helps forecast infra needs.
7. **Tiered / Overages Mix % (TIER)**\
   \&#xNAN;*EN:* Share of revenue from higher tiers or overage charges vs base rate.\
   \&#xNAN;*Pseudo:* `Revenue_from_Tiers_>1 / Total_Usage_Revenue * 100`\
   \&#xNAN;*Why:* Shows success of pricing design in capturing heavy users’ value.
8. **Unit Cost (Variable)**\
   \&#xNAN;*EN:* Direct variable cost per unit (bandwidth, compute, fuel).\
   \&#xNAN;*Pseudo:* `Variable_Costs / Total_Units`\
   \&#xNAN;*Why:* Drives GMU; optimizing infra/vendor deals pushes PPUU up.
9. **New Active Users (ACQ)**\
   \&#xNAN;*EN:* Fresh customers who generated usage this period.\
   \&#xNAN;*Pseudo:* `COUNT(users with first_usage_date in period)`\
   \&#xNAN;*Why:* Pipeline for future volume; complements expansion of existing users.
10. **Expansion Usage % (EXPu)**\
    \&#xNAN;*EN:* Incremental usage from existing users vs last period.\
    \&#xNAN;*Pseudo:* `(Usage_existing_t - Usage_existing_{t-1}) / Usage_existing_{t-1} * 100`\
    \&#xNAN;*Why:* Land-and-expand health metric in usage world.
11. **Growth Efficiency Factor (GEF)** *EN:* How efficiently you convert capacity (supply) growth into profitable usage growth. *Pseudo:* `GEF = (Demand_Growth % / Capacity_Growth %) * (GMU / 100)` *Why:* If you scale infra faster than demand or without keeping margins, you destroy PPUU. GEF >1 means you’re growing demand faster than capacity (or holding margin so growth is efficient). *Benchmark idea:* Cloud/API teams often target **GEF ≥ 1.2** in growth phases; utilities, constrained by regulation, hover near 1 (capacity expansions match load forecasts).
12. **Demand (Usage) Growth % (DG)** *EN:* Year-over-year change in total billed usage units. *Pseudo:* `(Usage_t − Usage_{t−1}) / Usage_{t−1} * 100` *Why:* Core top-line driver in metered models; faster than price hikes, usage growth proves product value. *Benchmark idea:* Early-stage APIs: **30–50%+ YoY**; mature utilities: **1–2% YoY**.
13. **Capacity Growth % (CG)** *EN:* YoY change in maximum deliverable capacity (compute throughput, MW, TPS). *Pseudo:* `(Cap_t − Cap_{t−1}) / Cap_{t−1} * 100` *Why:* Overbuild wastes capex; underbuild throttles revenue (throttling, outages). *Benchmark idea:* hyperscalers expand capacity \~in line with forecasted demand + buffer (e.g., **20–40% YoY** during high-growth years).
14. **Capacity Utilization % (CI)** *EN:* Average share of provisioned capacity actually used (often measured at peak window or averaged). *Pseudo:* `Avg_Usage / Provisioned_Capacity * 100` *Why:* Direct efficiency metric—too low means idle assets, too high means no headroom for spikes. *Benchmark idea:* Power plants target **\~60–80%** load factors; cloud infra teams often aim **40–60% average** to leave burst room.
15. **Provisioned Capacity (PC)** *EN:* The maximum sustained throughput you can deliver (e.g., kWh/day, requests/sec). *Pseudo:* `Cap = Σ(node_capacity)` (or grid MW installed) *Why:* Sets the ceiling; also denominator for utilization. Needed for capex planning.
16. **Peak Usage / Peak Load (PU)** *EN:* Highest instantaneous (or short-window) usage observed in the period. *Pseudo:* `max(usage_rate_t)` over period

    *Why:* Determines required headroom and auto-scaling needs; drives worst-case cost.

    *Benchmark:* Peak multiples of 1.5–3× average are common; extreme bursty APIs can see 10×.
17. **Peak-to-Average Ratio (PAR)** *EN:* Ratio of peak load to average load. *Pseudo:* `Peak_Usage / Average_Usage`

    *Why:* Quantifies burstiness; high PAR stresses infra cost and pricing design (need overage/tier pricing).

    *Benchmark:* Utilities PAR \~1.3–1.6; consumer APIs (chat/LLM) can be **>3–5** during viral events.
18. **Capex per Unit Capacity (CAPU)** *EN:* Capital required to add one unit of capacity (e.g., $ per kW, $ per 1k TPS). *Pseudo:* `Capex_added / Capacity_added`

    *Why:* Guides ROI on expansion; lower CAPU means cheaper scaling.

    *Benchmark:* Data center build costs ≈ **$7–12M per MW**; GPU clusters vary wildly but trend \~$25–40 per deployable TFLOP.
19. **Auto-Scaling Latency (ASL)** *EN:* Time it takes to provision additional capacity after demand spike. *Pseudo:* `t(scale_complete) - t(threshold_trigger)`

    *Why:* Slow scaling means you must over-provision; fast scaling lets you run lean.

    *Benchmark:* Best cloud-native infra targets **seconds–minutes**; regulated utilities can’t “auto-scale,” they plan years ahead.
20. **Headroom % (HR)** *EN:* Buffer capacity above expected peak. *Pseudo:* `(Provisioned_Capacity - Expected_Peak) / Provisioned_Capacity * 100`

    *Why:* Prevents outages and throttling; too much wastes capital.

    *Benchmark:* Cloud SREs often keep **20–30%** headroom; grid operators maintain N-1 redundancy (varies but \~15–25% capacity reserve).
21. **Usage Concentration Risk %** *(auxiliary)*\
    \&#xNAN;*EN:* % of total usage (or revenue) coming from top X customers.\
    \&#xNAN;*Pseudo:* `Usage_top_X / Total_Usage * 100`\
    \&#xNAN;*Why:* Whales are great—until they churn. Tracks fragility of revenue base.\
    \&#xNAN;*Benchmark:* Aim <30% from top 5; many infra/API firms start >50% and work it down over time.


# Licensing Intellectual Property

**Definition (short).** You monetize intellectual property—patents, code, characters, brands—by letting others use it under contract. Cash arrives as upfront license fees and ongoing royalties, typically a percentage of the licensee’s sales or a fixed dollar amount per unit.

**Recent example.** Qualcomm’s 5G handset portfolio license charges [5% of the net selling price (capped around $20 per 5G phone)](https://www.qualcomm.com/content/dam/qcomm-martech/dm-assets/documents/qualcomm-5g-handset-licensing-program.pdf?utm_source=chatgpt.com), generating more than $6 billion of royalties in FY2024.

**Historical example.** Bell Labs licensed the transistor in the 1950s, kickstarting the modern electronics industry, just as earlier inventors like Elias Howe lived off sewing-machine patent royalties.

**Another current datapoint.** ARM [reported $2.2 billion in royalty revenue in FY2022](https://stockdividendscreener.com/technology/arm-holdings-revenue-breakdown-by-segment/?utm_source=chatgpt.com), with its CPU IP inside roughly 95% of smartphones. **Consumer IP illustration.** Disney’s Consumer Products & Licensing is only about 5% of revenue but \~13% of operating income, showing how lucrative high-margin licensing can be.

<figure><img src="/files/1qgcCl0tMvDP5xV6Y2Pf" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

1. **Royalty Profit Rate % (RPR).** This is the percentage of royalty and license revenue that remains after you pay for IP enforcement and for creating or refreshing the IP. *Pseudo:* `(Royalty_Rev − Legal_IP_Costs − IP_R&D) / Royalty_Rev * 100` *Why it matters:* Licensing looks wonderfully high-margin until you subtract lawsuit costs and ongoing invention/content spend; this KPI tells you if the model truly throws off cash. *Benchmark:* Pure licensors routinely post 80–90% gross margins, but net after legal and refresh costs often lands in the 40–60% range.
2. **Royalty / License Revenue (RR).** This is the total dollars from ongoing royalties plus any upfront or milestone license fees. *Pseudo:* `Σ(licensee_sales × rate) + upfront_fees` *Why it matters:* It is your top line; without it none of the downstream percentages matter. *Benchmark:* ARM booked $2.2 billion in royalties in FY 22; Qualcomm exceeded $6 billion in FY 24.
3. **Gross Margin on IP Revenue % (GMIP).** This measures how much of RR you keep before legal and R\&D, i.e., after only direct delivery costs. *Pseudo:* `(RR − Direct_Delivery_Costs) / RR * 100` *Why it matters:* It shows how “software-like” your margin structure is at the top; any drop suggests expensive third-party data, rev-share pass-throughs, or inefficient operations. *Benchmark:* Many licensors report >85% gross margin on IP revenue.
4. **# of Licensees (LA).** This counts active licensing agreements or partners. *Pseudo:* `COUNT(active_licenses)` *Why it matters:* More licensees usually means broader adoption and less dependence on any single partner, although quality trumps raw count. *Benchmark:* ARM discloses 500+ active licensees, while large franchise systems run into the tens of thousands.
5. **Licensee Sales Base (LSB).** This is the aggregate sales or units sold by licensees on which your royalties are calculated. *Pseudo:* `Σ licensee_reported_sales_subject_to_royalty` *Why it matters:* Your revenue rides directly on LSB; tracking it helps forecast royalties and spot licensee underperformance early. *Benchmark:* Qualcomm’s LSB essentially equals the global smartphone market (\~1.2 billion units/year).
6. **Average Royalty Rate % or $/Unit (RRATE).** This is the effective percentage of licensee sales (or dollars per unit) you capture. *Pseudo:* `RR / LSB` *Why it matters:* It reflects how much value you’re extracting from the ecosystem; too low and you subsidize licensees, too high and they may seek workarounds. *Benchmark:* Patent deals commonly fall in 3–5%; character/trademark licensing 8–15%.
7. **Legal Enforcement Cost % of Rev (LEG).** This is the share of RR spent on litigation, monitoring, and IP enforcement. *Pseudo:* `Legal_IP_Spend / RR * 100` *Why it matters:* An unexpected lawsuit can wipe out a quarter’s profit; persistent high LEG means your moat is expensive to defend. *Benchmark:* Large licensors try to stay below 10%; heavy-litigation portfolios can spike above 20%.
8. **R\&D / Content Refresh % of Rev (RND).** This is how much of RR you reinvest in new patents or fresh content. *Pseudo:* `IP_R&D / RR * 100` *Why it matters:* Patents expire and characters age; sustained royalty streams require fresh IP. *Benchmark:* Qualcomm spends \~15–25% of total revenue on R\&D; Disney continually injects billions into new franchises.
9. **Top-5 Licensee Concentration % (CONC).** This is the percentage of RR coming from your five largest licensees. *Pseudo:* `RR_top5 / RR_total * 100` *Why it matters:* Over-reliance gives bargaining leverage to a few partners and raises revenue volatility risk. *Benchmark:* Aim for <50% from the top five.
10. **Licensee Sales Growth % (GROW).** This is the year-over-year growth in the LSB. *Pseudo:* `(LSB_t − LSB_{t-1}) / LSB_{t-1} * 100` *Why it matters:* If your licensees’ businesses stagnate, your royalties plateau unless you raise rates.
11. **% Deals with Minimum/Floor (MINF).** This is the share of contracts that include guaranteed minimum payments. *Pseudo:* `Deals_with_Minimums / Total_Deals * 100` *Why it matters:* Floors protect downside when a licensee’s sales underperform; too many minimums can scare small partners away. *Benchmark:* It is common to see >30% of deals in volatile sectors include minimums.
12. **IP Portfolio Quality Index (QUAL).** This is a weighted score of portfolio “strength” (e.g., share of standard-essential patents, citation counts, brand valuation). *Pseudo:* `w1*#SEPs + w2*Citation_Count + w3*Brand_Value ...` *Why it matters:* Higher-quality IP is easier to license, demands higher rates, and costs less to defend. *Benchmark:* Disney’s brand value exceeds $40 billion, while companies like IBM and Samsung each hold 100k+ active patents.


# Selling Service

**Definition (short).** You sell expert human time (and any pass-through materials) on an as-consumed basis—hours, days, or sprints. Revenue equals billable hours multiplied by realized rates; profitability depends on utilization, pricing discipline, and delivery efficiency.

**Recent example.** Accenture reported [$64.9 billion in FY 2024 revenue](https://newsroom.accenture.com/content/4q-full-fy24-earnings/accenture-reports-fourth-quarter-and-full-year-fiscal-2024-results.pdf?utm_source=chatgpt.com) with a GAAP operating margin of 14.8%.

**Historical example.** McKinsey (1926) and the Big Four accounting firms institutionalized the billable hour; law firms have tracked “billable vs. non-billable” time since the mid-20th century. **Rate reality check.** Elite U.S. law firms now bill [up to $3,000 per hour for senior partners](https://taxprof.typepad.com/taxprof_blog/2024/10/2025-hourly-billing-rates-for-elite-law-firms-3000-for-senior-partners-1000-for-first-year-associate.html), while first-year associates flirt with $1,000.

**Utilization benchmark.** SPI Research’s benchmarks put average billable utilization at \~67–68%.

<figure><img src="/files/rToIp5cYEwlReXDF3eVJ" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

1. **Profit per Productive Hour (PPH).** This is the operating profit you generate for every billable hour delivered. *Pseudo:* `PPH = (Revenue − Direct_Labor − Materials − Alloc_Opex) / Billable_Hours` *Why it matters:* It compresses the two levers you actually control—pricing and delivery efficiency—into one easy-to-compare dollar figure. Rising PPH means you are either charging more, staffing smarter, or avoiding scope creep. *Benchmark:* Mid-market consulting/IT firms often target $40–$80 profit per hour, whereas elite strategy and law firms can clear $150+.
2. **Billable Hours (BH).** This is the total number of hours billed to clients in the period. *Pseudo:* `Σ billable_hours` *Why it matters:* Hours are your “units sold.” Idle people are a pure cost. *Benchmark:* A typical consultant bills 1,400–1,800 hours/year, corresponding to \~70–85% utilization.
3. **Average Billing Rate $/hr (ABR).** This is the realized dollar amount per billable hour after discounts. *Pseudo:* `Services_Revenue / Billable_Hours` *Why it matters:* It reflects pricing power and role mix. If ABR drifts down, you are discounting or over-indexing on junior staff. *Benchmark:* Management consulting commonly $200–$400/hr; IT services $100–$250/hr; top law partners now up to $3,000/hr.
4. **Project Gross Margin % (GMPR).** This measures how much margin you make after direct labor and pass-through materials on a project. *Pseudo:* `(Revenue − Direct_Labor − Materials) / Revenue * 100` *Why it matters:* It is the clearest indicator of delivery efficiency and scoping accuracy. *Benchmark:* Healthy professional services firms run 30–50% project gross margin; top boutiques can exceed 50%.
5. **Utilization % (UTIL).** This is the percentage of available working hours that are billable. *Pseudo:* `Billable_Hours / Available_Hours * 100` *Why it matters:* Utilization is the core productivity lever. Too low means you are carrying bench cost; too high risks burnout and quality issues. *Benchmark:* Industry averages hover around 67–68%, while many firms target 75–80%.
6. **Billable Headcount (HC).** This is the number of employees whose time is sold to clients. *Pseudo:* `COUNT(billable_staff)` *Why it matters:* It defines capacity. Over-hiring without demand crushes utilization and PPH. *Benchmark:* Many firms keep 70–80% of total headcount billable (with the rest in support roles).
7. **Role / Rate Mix Index (MIX).** This is a weighted index showing the blend of partner/manager/analyst hours that drives ABR. *Pseudo:* `Σ(role_hours × role_rate) / Σ hours` *Why it matters:* A healthy pyramid keeps costs low and rates high; skewing too senior or too junior harms margin or quality. *Benchmark:* A common consulting pyramid is roughly 10–15% partners, 25–35% managers, remainder juniors.
8. **Discount / Scope Creep % (DISC).** This is the percentage of list billable value that you fail to collect because of discounts or out-of-scope work you did not bill. *Pseudo:* `(List_Value − Actual_Revenue) / List_Value * 100` *Why it matters:* It is a silent margin killer; many firms underestimate how much money leaks here. *Benchmark:* Elite firms keep leakage below 5%, while many agencies leak 10–15%.
9. **Direct Labor Cost % of Rev (DLC).** This is the proportion of project revenue consumed by salaries and benefits of billable staff. *Pseudo:* `Direct_Labor / Project_Revenue * 100` *Why it matters:* Labor is your biggest cost. Keeping DLC in line is essential to maintain GMPR. *Benchmark:* Targets are typically ≤40–50%; beyond 55% margins compress quickly.
10. **Materials / Pass-through % (MTR).** This is the share of revenue that is simply passed through (media spend, subcontractors, hardware) and carries little margin. *Pseudo:* `Pass_through / Revenue * 100` *Why it matters:* Pass-through inflates revenue but not profit; you need to separate it to read margins correctly. *Benchmark:* Media agencies can see >60% pass-through; classic consulting often <5%.
11. **Revenue per Employee (RPE).** This is total revenue divided by total headcount (billable + support). *Pseudo:* `Revenue / Total_Headcount` *Why it matters:* It is a macro productivity indicator; falling RPE often signals organizational bloat. *Benchmark:* Many PS firms target $200k–$300k per billable FTE, while top AmLaw 100 firms exceed $1 million per lawyer.


# Selling Outcomes

## Outcome / Performance-Based (Success / Contingency / Gainshare)

**Definition (short).** You only get paid (or get most of your pay) if a predefined outcome is achieved—e.g., a hire is made, a lawsuit is won, a deal closes, a KPI target is hit. You trade certainty for upside: lower or zero base fees, larger “success fees” when you deliver.

**Recent example.** Contingent recruiters typically charge [15 – 25 % of first-year salary](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com) only if the candidate is hired (e.g., $20 k on a $100 k role). Personal-injury lawyers in the U.S. commonly take [30 – 40 % of the award](https://www.americanbar.org/groups/legal_services/milvets/aba_home_front/information_center/working_with_lawyer/fees_and_expenses/?utm_source=chatgpt.com) on a “no win, no fee” basis. Investment banks collect “success fees” of [\~ 1 % on mid-market M\&A deals](https://firstpagesage.com/business/ma-advisory-fee-structure/?utm_source=chatgpt.com), falling as deal size rises. Performance marketing/affiliate programs pay only on conversion or sale.

**Historical example.** Real-estate brokers since the late 19th century have been paid a [5 – 6 % commission on sale price](https://www.investopedia.com/articles/active-trading/031215/how-real-estate-agent-and-broker-fees-work.asp?utm_source=chatgpt.com) (split between buyer and seller agents). Maritime “no cure, no pay” salvage contracts go back centuries: salvors only get paid if they save the ship or cargo ([Lloyd’s Open Form](https://en.wikipedia.org/wiki/Lloyd's_Open_Form?utm_source=chatgpt.com)).

<figure><img src="/files/yIQ5ZFwnPe1Q8Z7pNJiB" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

**Expected Value per Engagement (EVPE).**\
\&#xNAN;*Pseudo:* `EVPE = (Success_Rate * Fee_per_Success) − Cost_per_Opportunity`.\
\&#xNAN;*Why it matters:* A CEO needs to know if the portfolio of bets is positive EV; high EVPE means your pricing and screening compensate for failures.\
\&#xNAN;*Benchmark:* A healthy contingency recruiter might target [EVPE ≥ 2× direct delivery cost per search](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com), given \~25 % fill rates and 20 % fees.

**Success Rate % (SR).**\
\&#xNAN;*Pseudo:* `SR = Successful_Engagements / Total_Engagements * 100`.\
\&#xNAN;*Why it matters:* Low SR means wasted effort; high SR indicates strong screening or execution.\
\&#xNAN;*Benchmark:* Contingent recruiters report [20 – 30 % fill rates](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com); plaintiff lawyers often claim [60 – 90 % win/settle rates](https://www.americanbar.org/groups/legal_services/milvets/aba_home_front/information_center/working_with_lawyer/fees_and_expenses/?utm_source=chatgpt.com).

**Average Fee per Success (FPS).**\
\&#xNAN;*Pseudo:* `FPS = Total_Success_Fees / Successful_Engagements`.\
\&#xNAN;*Why it matters:* Together with SR it defines revenue potential; larger FPS can offset lower SR.\
\&#xNAN;*Benchmark:* Recruiting: [15 – 25 % of first-year salary](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com). M\&A advisory: [\~ 1 % on mid-sized deals](https://firstpagesage.com/business/ma-advisory-fee-structure/?utm_source=chatgpt.com). Legal contingency: [30 – 40 % of settlement](https://www.americanbar.org/groups/legal_services/milvets/aba_home_front/information_center/working_with_lawyer/fees_and_expenses/?utm_source=chatgpt.com).

**Cost per Opportunity Closed (COC).**\
\&#xNAN;*Pseudo:* `COC = (Sales + Delivery Costs for that Engagement)`.\
\&#xNAN;*Why it matters:* If COC creeps up, your EV per engagement shrinks; it’s the denominator for pricing.\
\&#xNAN;*Benchmark:* Top recruiters keep delivery cost [< 25 % of expected fee](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com).

**Total Opportunities Undertaken (OPP).**\
\&#xNAN;*Pseudo:* `COUNT(engagements_started)`.\
\&#xNAN;*Why it matters:* Volume drives total revenue potential but can dilute focus and SR if you take weak deals.\
\&#xNAN;*Benchmark:* A single recruiter handling [5 – 10 searches at once](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com) is typical.

**Average Cycle Time to Outcome (CT).**\
\&#xNAN;*Pseudo:* `mean(end_date − start_date)`.\
\&#xNAN;*Why it matters:* Long cycles tie up resources and delay cash. Faster cycles improve annualized EV.\
\&#xNAN;*Benchmark:* Recruiters aim [30 – 60 days to fill](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com).

**Client Value Delivered $ (VAL).**\
\&#xNAN;*Pseudo:* `Σ client_value_metric`.\
\&#xNAN;*Why it matters:* Validates fee size and proves ROI; helps negotiate better splits.\
\&#xNAN;*Benchmark:* Gainshare consulting commonly takes [10 – 30 % of measured savings](https://firstpagesage.com/business/ma-advisory-fee-structure/?utm_source=chatgpt.com).

**Provider % of Value (Split %) (SPLIT).**\
\&#xNAN;*Pseudo:* `Fee_per_Success / Client_Value * 100`.\
\&#xNAN;*Why it matters:* Ensures pricing aligns with delivered value; too low = leaving money on table; too high = client pushback.\
\&#xNAN;*Benchmark:* Typical splits: [10 – 30 % for gainshare consulting](https://firstpagesage.com/business/ma-advisory-fee-structure/?utm_source=chatgpt.com), [\~ 5 – 6 % realtor commission](https://www.investopedia.com/articles/active-trading/031215/how-real-estate-agent-and-broker-fees-work.asp?utm_source=chatgpt.com).

**Acquisition & Delivery Cost / Opportunity (ACQCost).**\
\&#xNAN;*Pseudo:* `Total_S&M + Delivery_Pre-fee / Opportunities`.\
\&#xNAN;*Why it matters:* This feeds the EVPE equation; if ACQCost rises faster than FPS or SR, your economics deteriorate.\
\&#xNAN;*Benchmark:* Growth-stage firms watch that ACQCost stays [< 50 % of EVPE](https://www.sage.com/en-us/-/media/files/sagedotcom/master/gated-assets/intacct/whitepapers/2023-professional-services-maturity-benchmark.pdf?utm_source=chatgpt.com).

**Risk Screening Score (RISK).**\
\&#xNAN;*Pseudo:* `score = f(evidence_strength, client_commitment, precedent, achievable_value)`\
\&#xNAN;*Why it matters:* Good screening improves SR and EVPE; bad screening floods pipeline with losers.\
\&#xNAN;*Benchmark:* Top contingency firms accept only [< 10 % of inbound PI cases](https://www.americanbar.org/groups/legal_services/milvets/aba_home_front/information_center/working_with_lawyer/fees_and_expenses/?utm_source=chatgpt.com).

**Qualified Pipeline % (PIPE).**\
\&#xNAN;*Pseudo:* `Qualified_Opps / Total_Opps * 100`.\
\&#xNAN;*Why it matters:* Focused pipelines save cost and maintain SR.\
\&#xNAN;*Benchmark:* Many high-end firms keep [\~ 30 – 50 % of inbound leads](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com).

**Work-in-Progress Load (WIP).**\
\&#xNAN;*Pseudo:* `COUNT(active_engagements)` or `Σ hours_in_progress`.\
\&#xNAN;*Why it matters:* WIP must match team capacity; overload reduces SR and elongates CT.\
\&#xNAN;*Benchmark:* Recruiters handle [5 – 10 live searches](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com).

**Client ROI Multiple (ROIc).**\
\&#xNAN;*Pseudo:* `Client_Value / Fee_Paid`.\
\&#xNAN;*Why it matters:* A solid ROI story justifies your fee structure and helps renewals/referrals.\
\&#xNAN;*Benchmark:* Gainshare deals often promise [3 – 10× client ROI](https://firstpagesage.com/business/ma-advisory-fee-structure/?utm_source=chatgpt.com).

**Cap / Floor Clauses % (CAP).**\
\&#xNAN;*Pseudo:* `Deals_with_cap_or_floor / Total_Deals * 100`.\
\&#xNAN;*Why it matters:* Caps can limit upside; floors protect downside—balance both.\
\&#xNAN;*Benchmark:* Recruiters often have [minimum flat fees (\~$10 k)](https://www.recruiterslineup.com/contingency-recruiting-fee-structure/?utm_source=chatgpt.com).

**Top-5 Clients Revenue Concentration % (CONC) (aux).**\
\&#xNAN;*Pseudo:* `Rev_top5 / Total_Rev * 100`.\
\&#xNAN;*Why it matters:* A whale leaving can crush EVPE and SR.\
\&#xNAN;*Benchmark:* Aim for [< 50 %](https://www.sage.com/en-us/-/media/files/sagedotcom/master/gated-assets/intacct/whitepapers/2023-professional-services-maturity-benchmark.pdf?utm_source=chatgpt.com).


# Taking a cut (Marketplaces)

**Definition (short).** You connect two (or more) sides of a transaction (buyers–sellers, riders–drivers, hosts–guests) and charge a commission or fee. You rarely own the underlying good/service; your economics are driven by **GMV volume × take rate**, minus operating costs.

**Recent example.** Airbnb’s effective take rate was [\~ 13 – 15 % on $60 B+ GBV in 2023](https://fourweekmba.com/airbnb-take-rate/?utm_source=chatgpt.com). Uber takes [\~ 20 – 28 % of each fare](https://stockanalysis.com/stocks/uber/metrics/?utm_source=chatgpt.com) after incentives. Amazon Marketplace charges [8 – 15 % category commissions](https://productscope.ai/blog/amazon-referral-fee/?utm_source=chatgpt.com); eBay averages \~ 10 %.

**Historical example.** Sotheby’s (1744) and Christie’s (1766) have always taken [auction commissions](https://www.investopedia.com/articles/active-trading/031215/how-real-estate-agent-and-broker-fees-work.asp?utm_source=chatgpt.com). Stockbrokers since the 17th century charged per trade, and newspaper classifieds were proto-marketplaces taking small fees to connect local buyers and sellers.

<figure><img src="/files/5j46BbYD8ERESLEcFNJo" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

**Net Commission Profit Growth % (NCPG).**\
\&#xNAN;*Pseudo:* `((Commission_Rev - Operating_Costs)_t - (Commission_Rev - Operating_Costs)_{t-1}) / (Commission_Rev - Operating_Costs)_{t-1} * 100`.\
\&#xNAN;*Why it matters:* GMV without monetization, or take rate without efficiency, equals vanity. NCPG unifies the three real levers: volume (GMV), monetization (take rate), and cost discipline.\
\&#xNAN;*Benchmark:* Mature marketplaces target [double-digit profit growth](https://www.sec.gov/Archives/edgar/data/1559720/000155972024000006/abnb-20231231.htm?utm_source=chatgpt.com).

**Gross Merchandise / Transaction Value (GMV).**\
\&#xNAN;*Pseudo:* `Σ transaction_amounts`.\
\&#xNAN;*Why it matters:* Commission dollars usually scale linearly with GMV. Investors use GMV to size the marketplace even before profits.\
\&#xNAN;*Benchmark:* Airbnb GBV [$63 B (2022)](https://www.sec.gov/Archives/edgar/data/1559720/000155972024000006/abnb-20231231.htm?utm_source=chatgpt.com).

**Take Rate % (TR).**\
\&#xNAN;*Pseudo:* `TR = Commission_Revenue / GMV * 100`.\
\&#xNAN;*Why it matters:* It shows value capture and pricing power; raising TR without hurting liquidity is gold.\
\&#xNAN;*Benchmark:* Product marketplaces [5 – 15 %](https://productscope.ai/blog/amazon-referral-fee/?utm_source=chatgpt.com); service marketplaces 15 – 30 %.

**Operating Cost % of Commission Rev (OPEX).**\
\&#xNAN;*Pseudo:* `OPEX / Commission_Revenue * 100`.\
\&#xNAN;*Why it matters:* High opex eats into take rate gains; scale should drive this down.\
\&#xNAN;*Benchmark:* Post-scale marketplaces target [< 50 % opex/commission rev](https://www.sec.gov/Archives/edgar/data/1559720/000155972024000006/abnb-20231231.htm?utm_source=chatgpt.com).

**Number of Transactions (NT).**\
\&#xNAN;*Pseudo:* `COUNT(txn_id)`.\
\&#xNAN;*Why it matters:* Frequency matters for habit and network effects; many small orders can equal one large one in GMV, but frequency builds stickiness.\
\&#xNAN;*Benchmark:* DoorDash did [512 million orders in Q1 2023](https://ir.doordash.com/news/news-details/2023/DoorDash-Releases-First-Quarter-2023-Financial-Results/default.aspx?utm_source=chatgpt.com).

**Average Transaction Value $ (ATV).**\
\&#xNAN;*Pseudo:* `ATV = GMV / NT`.\
\&#xNAN;*Why it matters:* Shifts in product mix or user behavior drive ATV; high ATV can boost revenue per txn but often comes with lower frequency.\
\&#xNAN;*Benchmark:* E-commerce AOV [$50 – $100](https://www.irpcommerce.com/en/us/ecommercemarketdata.aspx?utm_source=chatgpt.com).

**Other Fees % of GMV (OTHERF).**\
\&#xNAN;*Pseudo:* `Other_Fee_Revenue / GMV * 100`.\
\&#xNAN;*Why it matters:* Diversifies income and effectively raises take rate without raising headline commission.\
\&#xNAN;*Benchmark:* Etsy charges [listing + promoted listings fees](https://investors.etsy.com/overview/key-figures/default.aspx?utm_source=chatgpt.com) adding a few extra %.

**Mix of Commission vs Subscriptions % (MIXT).**\
\&#xNAN;*Pseudo:* `Subscription_Revenue / Total_Revenue * 100`.\
\&#xNAN;*Why it matters:* Subscriptions stabilize revenue and raise LTV; but too much may deter small sellers.\
\&#xNAN;*Benchmark:* Some B2B marketplaces get [> 20 % of revenue from subscriptions](https://investors.etsy.com/overview/key-figures/default.aspx?utm_source=chatgpt.com).

**Active Buyers (ACTB).**\
\&#xNAN;*Pseudo:* `COUNT(DISTINCT buyer_id WHERE txn>0)`.\
\&#xNAN;*Why it matters:* Demand-side depth; more buyers improve liquidity, justify higher TR.\
\&#xNAN;*Benchmark:* Etsy had [95.1 M active buyers (2022)](https://app.stocklight.com/stocks/us/nasdaq-etsy/etsy/annual-reports/nasdaq-etsy-2023-10K-23655284.pdf?utm_source=chatgpt.com).

**Active Sellers (ACTS).**\
\&#xNAN;*Pseudo:* `COUNT(DISTINCT seller_id WHERE txn>0)`.\
\&#xNAN;*Why it matters:* Supply depth; too few sellers causes stockouts/price spikes; too many with no sales drives churn.\
\&#xNAN;*Benchmark:* Etsy reported [7.5 M active sellers (2022)](https://app.stocklight.com/stocks/us/nasdaq-etsy/etsy/annual-reports/nasdaq-etsy-2023-10K-23655284.pdf?utm_source=chatgpt.com).

**Buyer Conversion Rate % (CONV).**\
\&#xNAN;*Pseudo:* `Purchases / Qualified_Sessions * 100`.\
\&#xNAN;*Why it matters:* It signals liquidity and UX quality; higher conversion usually lifts both GMV and TR.\
\&#xNAN;*Benchmark:* General e-commerce sees [2 – 3 % conversion rates](https://www.invespcro.com/cro/conversion-rate-by-industry/?utm_source=chatgpt.com).

**Purchase Frequency / Buyer (LFQ).**\
\&#xNAN;*Pseudo:* `NT / Active_Buyers`.\
\&#xNAN;*Why it matters:* Frequency drives lifetime value and defensibility.\
\&#xNAN;*Benchmark:* Top “habit” marketplaces push [> 5 – 10 orders per buyer per year](https://ir.doordash.com/news/news-details/2023/DoorDash-Releases-First-Quarter-2023-Financial-Results/default.aspx?utm_source=chatgpt.com).

**Supply Health / Listings Liquidity % (SUPH).**\
\&#xNAN;*Pseudo:* `Listings_sold / Listings_posted_in_window * 100`.\
\&#xNAN;*Why it matters:* If sellers don’t sell, they churn; liquidity is the soul of a marketplace.\
\&#xNAN;*Benchmark:* Healthy resale marketplaces target [> 50 % sell-through over 60 – 90 days](https://ir.doordash.com/news/news-details/2023/DoorDash-Releases-First-Quarter-2023-Financial-Results/default.aspx?utm_source=chatgpt.com).

**Disintermediation / Leakage Rate % (RISK) (aux).**\
\&#xNAN;*Pseudo:* `Off_platform_transactions_est / Total_matches_est * 100`.\
\&#xNAN;*Why it matters:* High leakage erodes take rate and weakens the moat.\
\&#xNAN;*Benchmark:* Service marketplaces often battle [> 20 % leakage early on](https://investors.etsy.com/overview/key-figures/default.aspx?utm_source=chatgpt.com).


# AirBnB

**Definition (short).** You connect two sides (guests–hosts) and monetize via service fees. Core economics = **Gross Booking Value (GBV) × implied take rate − operating costs**. GBV is the dollar value of bookings (host earnings + fees + taxes, net of cancellations). ([SEC](https://www.sec.gov/Archives/edgar/data/1559720/000119312522138654/d711122dex991.htm?utm_source=chatgpt.com))

**Recent example.** Q1’25: **GBV $24.5B**, **Nights & Experiences 143.1M**, **ADR $171**, **Revenue $2.3B**, **Net income $154M (7% margin)**, **Adj. EBITDA $417M (18%)**, **TTM FCF $4.4B (39%)**. Implied take rate = **\~9.3%** for the quarter.

**Historical example.** Auction houses (Sotheby’s 1744, Christie’s 1766) long charged commissions to match buyers and sellers—Airbnb is the digital analog for lodging. ([Airbnb Newsroom](https://news.airbnb.com/airbnb-q4-2023-and-full-year-financial-results/?utm_source=chatgpt.com))

<figure><img src="/files/spVqtc3RUehmDwnw7SNO" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

**Gross Booking Value (GBV).**\
\&#xNAN;*Pseudo:* `Σ booking_amounts (host earnings + fees + taxes − cancels)`\
\&#xNAN;*Why it matters:* Top-of-funnel $$ volume; everything downstream scales from here. ([SEC](https://www.sec.gov/Archives/edgar/data/1559720/000119312522138654/d711122dex991.htm?utm_source=chatgpt.com))\
\&#xNAN;*Benchmark:* $24.5B in Q1’25; $82B+ in 2024. ([FourWeekMBA](https://fourweekmba.com/airbnb-gross-booking-value/?utm_source=chatgpt.com))

**Nights & Experiences Booked (NEB).**\
\&#xNAN;*Pseudo:* `COUNT(booked_units_net)`\
\&#xNAN;*Why it matters:* Core demand/usage driver of GBV.\
\&#xNAN;*Benchmark:* 143.1M in Q1’25 (+8% YoY).

**Average Daily Rate (ADR).**\
\&#xNAN;*Pseudo:* `ADR = GBV / Nights & Experiences Booked`\
\&#xNAN;*Why it matters:* Price/mix lever (geo, FX, LOS, product type).\
\&#xNAN;*Benchmark:* $171 in Q1’25 (−1% YoY; +1% ex-FX).

**Implied Take Rate % (TR).**\
\&#xNAN;*Pseudo:* `Revenue / GBV * 100`\
\&#xNAN;*Why it matters:* Monetization efficiency—how much value Airbnb captures. ([SEC](https://www.sec.gov/Archives/edgar/data/1559720/000119312522138654/d711122dex991.htm?utm_source=chatgpt.com))\
\&#xNAN;*Benchmark:* 9.3% in Q1’25; \~13.6% for FY 2024. ([FourWeekMBA](https://fourweekmba.com/how-much-does-airbnb-take/?utm_source=chatgpt.com))

**Revenue.**\
\&#xNAN;*Pseudo:* `GBV × TR`\
\&#xNAN;*Why it matters:* Primary operating outcome they steer to before costs.\
\&#xNAN;*Benchmark:* $2.3B in Q1’25 (+6% YoY; +8% ex-FX).

**Adjusted EBITDA (and Margin).**\
\&#xNAN;*Pseudo:* `Revenue − (Cash Opex + CoR) +/− adjustments`\
\&#xNAN;*Why it matters:* Management’s efficiency lens (ex some non-cash/one-offs).\
\&#xNAN;*Benchmark:* $417M (18% margin) in Q1’25.

**Net Income (and Margin).**\
\&#xNAN;*Pseudo:* `GAAP bottom line`\
\&#xNAN;*Why it matters:* Ultimate profitability after SBC, taxes, interest.\
\&#xNAN;*Benchmark:* $154M (7% margin) in Q1’25.

**Free Cash Flow (TTM).**\
\&#xNAN;*Pseudo:* `Operating Cash Flow − Capex`\
\&#xNAN;*Why it matters:* Real cash to fund buybacks/new bets; validates model durability.\
\&#xNAN;*Benchmark:* $4.4B TTM (39% margin).

**Active Listings / Supply Health (contextual driver).**\
\&#xNAN;*Pseudo:* `Distinct live listings; removals of low-quality supply`\
\&#xNAN;*Why it matters:* Liquidity & match quality sustain NEB and ADR. ([Airbnb Newsroom](https://news.airbnb.com/airbnb-q4-2023-and-full-year-financial-results/?utm_source=chatgpt.com))\
\&#xNAN;*Benchmark:* 7.7M active listings at YE 2023; 450k low-quality removed since 2023. ([Airbnb Newsroom](https://news.airbnb.com/airbnb-q4-2023-and-full-year-financial-results/?utm_source=chatgpt.com))

***

Let me know if you want this dropped into slides or a SQL/Python model.


# Selling Ads

**Definition (short).** You give users content or utility (often free or subsidized) and monetize by selling their attention and/or data to advertisers or sponsors. Revenue scales with audience size, engagement (impressions), and the price per impression/click/action.

**Recent example.** Meta booked **$164.5 billion in 2024 revenue, \~97% from ads** across Facebook, Instagram, and now Threads. [Meta booked $164.5 billion in 2024 revenue, \~97% from ads](https://investor.atmeta.com/investor-news/press-release-details/2025/Meta-Reports-Fourth-Quarter-and-Full-Year-2024-Results/default.aspx?utm_source=chatgpt.com). YouTube ads brought Alphabet **$10.47 billion in Q4 2024 (up 13.8%)**, while Google overall still derived **\~76% of 2024 revenue from ads**. [Google overall still derived \~76% of 2024 revenue from ads](https://www.marketingdive.com/news/google-ad-revenue-growth-sluggish-ai/739274/?utm_source=chatgpt.com).\
**Historical example.** U.S. broadcast TV has been ad-funded since the 1950s; newspapers in the late 19th century (e.g., *NYT*) sold papers cheaply and made money on classifieds and display ads—pioneering CPM pricing long before digital.

<figure><img src="/files/9VsnMUu4603ALCsHAACD" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

**Net Ad Profit Growth % (NAPR).** Year-over-year growth of (Ad Revenue − Direct Ad Ops Costs).\
\&#xNAN;*Pseudo:* `((AdRev - AdOpsCost)_t - (AdRev - AdOpsCost)_{t-1}) / (AdRev - AdOpsCost)_{t-1} * 100`.\
\&#xNAN;*Why it matters:* It unifies scale (revenue), monetization efficiency (eCPM/fill), and cost control. CEOs don’t want just bigger top-line; they want profitable growth.\
\&#xNAN;*Benchmark:* Mature ad giants are posting **\~20%+ YoY ad revenue growth with expanding margins** post-2023 rebound—top decile for scale players. [20%+ YoY ad revenue growth with expanding margins](https://investor.atmeta.com/investor-news/press-release-details/2025/Meta-Reports-Fourth-Quarter-and-Full-Year-2024-Results/default.aspx?utm_source=chatgpt.com)

**Ad Revenue $ (AR).** Total dollars from advertising/sponsorship in the period.\
\&#xNAN;*Pseudo:* `Σ(ad_impressions * price_per_impression) + sponsorship_fees`.\
\&#xNAN;*Why it matters:* It’s the core top line for an ad business; almost every other KPI feeds into it.\
\&#xNAN;*Benchmark:* Meta ads **$160.6 B in 2024**; Alphabet ads **\~76% of $307 B revenue**; ByteDance surpassed Meta in Q1 2025 ad revenue (\~$43 B vs $42.3 B).

**Ad Margin % (AM).** Gross margin on ad revenue after traffic acquisition costs (TAC), content costs, sales ops.\
\&#xNAN;*Pseudo:* `(AdRev − AdDirectCosts) / AdRev * 100`.\
\&#xNAN;*Why it matters:* High-volume inventory can still be low-margin if TAC is high. Margin signals pricing power and efficiency.\
\&#xNAN;*Benchmark:* Large platforms often report **70–80%+ gross margin on ad revenue** (digital distribution is cheap once built).

**Total Ad Impressions (IMP).** Count of ads actually shown.\
\&#xNAN;*Pseudo:* `COUNT(ad_served_id)`.\
\&#xNAN;*Why it matters:* Impressions are your monetizable inventory. Inventory growth (with stable CPM) = revenue growth.\
\&#xNAN;*Benchmark:* Social feeds can serve **dozens of ads per user per day** (e.g., \~40–50 on Facebook).

**Effective CPM/CPC/RPM (eCPM).** Average price per 1,000 impressions (or per click).\
\&#xNAN;*Pseudo:* `eCPM = (AdRev / Impressions) * 1000`.\
\&#xNAN;*Why it matters:* Shows monetization quality. Better targeting/format raises eCPM.\
\&#xNAN;*Benchmark:* Facebook/Meta global CPM hovered around **$13–14 in late 2024**; YouTube CPM **$2–$15 depending on niche/geo**. [Facebook/Meta global CPM hovered around $13–14 in late 2024](https://www.barrons.com/articles/meta-threads-platform-ads-0d941270?utm_source=chatgpt.com)

**Fill Rate % (FILL).** Percentage of ad slots filled with paid ads.\
\&#xNAN;*Pseudo:* `Paid_Impressions / Total_Available_Slots * 100`.\
\&#xNAN;*Why it matters:* Low fill = lost revenue or weak demand. High fill supports stable CPMs.\
\&#xNAN;*Benchmark:* Top platforms approach **\~100% fill** via auction backfill; smaller publishers may sit at **50–70%**.

**Format Mix % (MIXF).** Share of revenue from each ad format (video, display, search, native).\
\&#xNAN;*Pseudo:* `Rev_format / AdRev * 100`.\
\&#xNAN;*Why it matters:* Video has higher CPM but fewer slots; mix optimization drives ARPU.\
\&#xNAN;*Benchmark:* YouTube video ads grew **13.8% YoY in Q4 2024** and now form a rising share of Alphabet’s ad mix.

**Audience Size (MAU/DAU/Reach) (AUD).** Unique users/viewers in a period.\
\&#xNAN;*Pseudo:* `COUNT(DISTINCT user_id within window)`.\
\&#xNAN;*Why it matters:* Advertisers pay for reach; network value scales with audience.\
\&#xNAN;*Benchmark:* Meta MAUs \~3 B; Threads **300 M MAU in 2025**; large TV events still reach tens of millions.

**DAU/MAU % (STICK).** Stickiness—how many monthly users come daily.\
\&#xNAN;*Pseudo:* `DAU / MAU * 100`.\
\&#xNAN;*Why it matters:* Higher stickiness = more daily impressions and dependable inventory.\
\&#xNAN;*Benchmark:* Facebook historically \~**66–70% DAU/MAU**.

**Engagement / Impressions per User (ENG).** Average ads served per active user.\
\&#xNAN;*Pseudo:* `IMP / Active_Users`.\
\&#xNAN;*Why it matters:* Deeper engagement = more sellable slots per user.\
\&#xNAN;*Benchmark:* Heavy social users see **1000+ ads/month**; typical site visitors see far fewer.

**Average Time Spent / Sessions (ATS).** Minutes per user per day or sessions/user.\
\&#xNAN;*Pseudo:* `Σ minutes_viewed / Active_Users`.\
\&#xNAN;*Why it matters:* Time is a proxy for potential ad load without harming UX.\
\&#xNAN;*Benchmark:* TikTok/YouTube top \~**40–60 min/day**; Facebook/Instagram \~33 min/day.

**Yield Management Uplift % (YIELD).** Incremental revenue from optimization (better targeting, floor prices, header bidding).\
\&#xNAN;*Pseudo:* `(Optimized_Rev − Baseline_Rev)/Baseline_Rev * 100`.\
\&#xNAN;*Why it matters:* Smart yield lifts revenue without more ads.\
\&#xNAN;*Benchmark:* Smart bidding/header bidding rollouts often add **5–15% yield**.

**Ad Load / UX Risk Index (RISK) (aux).** Internal index of ad density vs. user satisfaction.\
\&#xNAN;*Pseudo:* `ad_slots_per_session / UX_score`.\
\&#xNAN;*Why it matters:* Over-monetization kills engagement; this guardrail protects long-term NAPR.\
\&#xNAN;*Benchmark:* Facebook publicly capped feed ad load increases after reaching “meaningful saturation” (\~15% of feed items).


# Financial Services

**Definition (short).** You earn by pricing and absorbing financial risk: banks capture **interest spread** (lend high, borrow low), insurers keep **premium minus claims/expenses** (combined ratio <100%). Profitability rests on margin (spread/ratio), volume (assets/premiums), and risk (defaults/claims).

**Recent example.** U.S. banks’ average **net interest margin was 3.30% in 2023** (up 35 bps YoY). [Net interest margin was 3.30% in 2023](https://www.fdic.gov/analysis/quarterly-banking-profile/qbp/2023dec/qbp.pdf?utm_source=chatgpt.com). U.S. P\&C insurers’ combined ratio improved to **101.7% in 2023** (still an underwriting loss), then to **\~96.5% in 2024** (profit) as price hikes beat claims. [Combined ratio improved to 101.7% in 2023, then to \~96.5% in 2024](https://www.ft.com/content/74109807-82d8-4281-a842-7681af0366fa?utm_source=chatgpt.com).\
**Historical example.** Lloyd’s of London (1680s) pioneered marine underwriting; medieval moneylenders lived on interest spreads; 19th-century savings & loans profited on mortgage spreads; the “combined ratio” has been an insurance staple since early 20th century.

<figure><img src="/files/7swjmi4zepduweB5a5YJ" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

**Net Interest/Underwriting Spread Profit Growth % (NISG).** YoY growth of profit from core spread/underwriting before investment gains.\
\&#xNAN;*Pseudo:* `((Spread_Profit)_t - (Spread_Profit)_{t-1}) / (Spread_Profit)_{t-1} * 100`.\
\&#xNAN;*Why it matters:* Captures whether you’re expanding the core engine (margin×volume) after losses/claims.\
\&#xNAN;*Benchmark:* Post-rate hikes, many U.S. banks grew spread profit double digits in 2023; P\&C insurers returned to underwriting profit in 2024 (combined \~96.5%).

**Net Interest Margin % (NIM).** (Banks) Interest income minus interest expense, over average earning assets.\
\&#xNAN;*Pseudo:* `(IntInc − IntExp) / Avg_Earning_Assets * 100`.\
\*Why it matters

:\* It’s the core spread metric; a few bps swing moves billions for big banks.\
\&#xNAN;*Benchmark:* U.S. banks **3.30% in 2023**; large money-center banks 2.6–3.2%; credit cards lenders 5–7%.

**Combined Ratio % (CR).** (Insurers) Loss ratio + expense ratio; <100% means underwriting profit.\
\&#xNAN;*Pseudo:* `(Claims/Premiums + Expenses/Premiums) * 100`.\
\&#xNAN;*Why it matters:* The single best underwriting KPI—tells if pricing outweighs claims+costs.\
\&#xNAN;*Benchmark:* U.S. P\&C **101.7% in 2023** (loss), **96.5% in 2024** (profit). Top performers target **<90–95%**.

**Volume of Earning Assets / Premiums (VOL).** Loans outstanding or gross written premiums.\
\&#xNAN;*Pseudo:* `Σ(loans) or Σ(premiums)`.\
\&#xNAN;*Why it matters:* Spread × volume = dollars. Scale matters, but not at the expense of risk.\
\&#xNAN;*Benchmark:* Healthy banks target **5–10% loan growth**; P\&C premium growth mid-single digits, spiking during “hard markets”.

**Cost of Funds % (COF).** Average interest rate paid on deposits/borrowings.\
\&#xNAN;*Pseudo:* `IntExp / Avg_Funding * 100`.\
\&#xNAN;*Why it matters:* Lower COF widens NIM without raising borrower rates.\
\&#xNAN;*Benchmark:* Large banks enjoy low-cost checking (near 0% for years), now rising but still below wholesale rates.

**Average Yield on Assets % (YLD).** Interest/ premium yield on loans/investments.\
\&#xNAN;*Pseudo:* `IntInc / Avg_Earning_Assets * 100`.\
\&#xNAN;*Why it matters:* Determines top side of NIM; asset mix and rate environment drive it.\
\&#xNAN;*Benchmark:* With 2023–24 rate hikes, banks’ asset yields rose several hundred bps; insurers’ bond portfolio new-money yields moved toward **4–5%**.

**Loss Ratio % (LR).** Claims paid ÷ premiums (insurance).\
\&#xNAN;*Pseudo:* `Claims / Premiums * 100`.\
\&#xNAN;*Why it matters:* Core risk pricing—too high implies underpricing or bad luck (cats).\
\&#xNAN;*Benchmark:* Personal lines saw LR >70% in 2023; after rate hikes, LR fell, driving combined ratio down to 96.5%.

**Expense Ratio % (ER).** Operating expenses ÷ premiums.\
\&#xNAN;*Pseudo:* `Underwriting_Expenses / Premiums * 100`.\
\&#xNAN;*Why it matters:* Operational efficiency lever; trimming ER improves CR directly.\
\&#xNAN;*Benchmark:* Many P\&C insurers run **\~25–30%** ER; best-in-class trims to low 20s.

**Loan/Premium Growth % (GROW).** Period-over-period growth in loans or premiums.\
\&#xNAN;*Pseudo:* `(ThisPeriod − LastPeriod) / LastPeriod * 100`.\
\&#xNAN;*Why it matters:* Indicates demand, share gain, and pricing power.\
\&#xNAN;*Benchmark:* Banks: **mid-single digits normal**, 10%+ aggressive. Insurers: **3–5% typical**, >10% in hard market cycles.

**Loan-to-Deposit Ratio % / Policy Retention % (LDR).** Banks: loans ÷ deposits; Insurers: % of policies renewed.\
\&#xNAN;*Pseudo:* `Loans/Deposits * 100` or `Policies_Renewed/Policies_Up * 100`.\
\&#xNAN;*Why it matters:* Funding stability (banks) and sticky business (insurers).\
\&#xNAN;*Benchmark:* Banks target **80–90% L/D**; Insurers retention **>85–90%** in personal lines is strong.

**Non-Performing Loan % / Charge-off % (NPL).** Share of loans in default or written off.\
\&#xNAN;*Pseudo:* `NPL / Total_Loans * 100`.\
\&#xNAN;*Why it matters:* Credit quality; high NPL erodes spread via provisions.\
\&#xNAN;*Benchmark:* U.S. banks’ NPL often **<2%** in good times; credit cards charge-offs **\~3%**.

**Catastrophe Loss % of Premiums (CAT).** P\&C cat claims as a % of earned premiums.\
\&#xNAN;*Pseudo:* `CatLosses / Premiums * 100`.\
\&#xNAN;*Why it matters:* Volatility driver; high CAT years wreck CR.\
\&#xNAN;*Benchmark:* 2024 saw $320 B global disaster losses; U.S. insurers still hit **96.5% CR** thanks to pricing.

**Investment Yield % (INVY) (aux).** Return on invested assets/float.\
\&#xNAN;*Pseudo:* `Investment_Income / Invested_Assets * 100`.\
\&#xNAN;*Why it matters:* Insurers rely on investment income to offset underwriting cycles; banks too on securities portfolios.\
\&#xNAN;*Benchmark:* With rates up, insurers’ new money yields **\~4–5%**; bonds dominated portfolios.


# Franchising

**Definition (short).** You let others (franchisees) operate under your brand and system. You earn initial fees plus **ongoing royalties (usually % of gross sales)** and often ad fund contributions and rent. You supply brand, playbooks, supply chain leverage; franchisees supply capital and operations.

**Recent example.** U.S. fast-food franchises typically pay **4–9% royalty on gross sales plus 2–4% advertising fees**. [4–9% royalty on gross sales plus 2–4% advertising fees](https://franzy.com/blog/average-franchise-royalty-fee/?utm_source=chatgpt.com). McDonald’s charges **\~4% royalty plus rent (\~8–12%)**, per its 2024 Franchise Disclosure Document. [McDonald’s charges \~4% royalty plus rent (\~8–12%)](https://www.franchimp.com/?f=107211_2024.pdf\&page=pdf\&utm_source=chatgpt.com).\
**Historical example.** Singer Sewing Machine (1850s) is cited as an early franchisor; modern franchising exploded post-WWII with chains like McDonald’s (1955), Holiday Inn, and Dunkin’.

<figure><img src="/files/bSEPGQrH36fY55bIYn6y" alt=""><figcaption></figcaption></figure>

#### KPI Definitions

**Net Franchise Royalty/Profit Growth % (NFRG).** YoY growth of (royalties + rents + fees − franchisor support costs).\
\&#xNAN;*Pseudo:* `((FranchiseProfit)_t - (FranchiseProfit)_{t-1}) / (FranchiseProfit)_{t-1} * 100`.\
\&#xNAN;*Why it matters:* Shows if expanding units and improving unit economics actually translate into franchisor profit.\
\&#xNAN;*Benchmark:* Mature systems still target **mid- to high-single-digit profit growth** from royalties/rents annually. [Mid- to high-single-digit profit growth from royalties/rents annually](https://www.franchisechatter.com/2024/10/13/fdd-talk-mcdonalds-franchise-costs-fees-average-revenues-and-or-profits-2024-review/?utm_source=chatgpt.com)

**Royalty & Fee Revenue $ (FREV).** Total ongoing royalties + initial fees + ad fund admin fees (if recognized).\
\&#xNAN;*Pseudo:* `Σ(royalties + initial_fees + other_recurring_fees)`.\
\&#xNAN;*Why it matters:* Core monetization stream; growth comes via more units and higher AUV.\
\&#xNAN;*Benchmark:* Royalties typically **4–9% of sales**; ad fees extra **2–4%**.

**Franchise Margin % (FMAR).** Gross margin on franchise operations (royalties & fees − franchise support costs).\
\&#xNAN;*Pseudo:* `(FRev − Support_Costs) / FRev * 100`.\
\&#xNAN;*Why it matters:* High-margin royalties can be eroded by heavy field support/legal costs.\
\&#xNAN;*Benchmark:* Asset-light franchisors often run **50–70%+ margins** on royalty revenue.

**Active Franchise Units (UNIT).** Number of open franchised outlets.\
\&#xNAN;*Pseudo:* `COUNT(franchise_locations_open)`.\
\&#xNAN;*Why it matters:* Scale equals revenue base; unit growth is the engine for future royalties.\
\&#xNAN;*Benchmark:* McDonald’s >**36k units worldwide**, Subway >**36k**, many top brands grow net units 2–4%/yr.

**Royalty % of Franchisee Sales (RPF).** Contracted royalty rate on gross sales.\
\&#xNAN;*Pseudo:* `RoyaltyPaid / Franchisee_Sales * 100`.\
\&#xNAN;*Why it matters:* Your “take rate”. Too high hurts franchisee economics; too low caps monetization.\
\&#xNAN;*Benchmark:* **4–9% typical**; McDonald’s \~4% plus rent; Chick-fil-A’s model is different (higher share of profit).

**New Units Opened (OPEN).** Count of openings in period (net of closures).\
\&#xNAN;*Pseudo:* `Openings − Closures`.\
\&#xNAN;*Why it matters:* Pipeline health and brand demand.\
\&#xNAN;*Benchmark:* Strong systems add **>3% net units/year**; weak ones shrink.

**Franchisee Retention % (RETEN).** Percentage of franchisees renewing at term or not selling back.\
\&#xNAN;*Pseudo:* `Renewed_Franchisees / Franchisees_Up_for_Renewal * 100`.\
\&#xNAN;*Why it matters:* Low churn signals healthy economics and satisfaction.\
\&#xNAN;*Benchmark:* Best systems keep **>90%** renewal; high churn flags trouble.

**Ad Fund % of Sales (ADVF).** Mandatory marketing fund contribution.\
\&#xNAN;*Pseudo:* `AdFee / Sales * 100`.\
\&#xNAN;*Why it matters:* Funds national marketing; too high squeezes operators.\
\&#xNAN;*Benchmark:* Often **2–4%**.

**Real Estate/Rent % of Sales (RENT).** Rent/franchise real estate markup as % of sales (for landlord-franchisors).\
\&#xNAN;*Pseudo:* `RentPaid_to_Franchisor / Sales * 100`.\
\&#xNAN;*Why it matters:* Major profit lever for brands that own/lease sites (e.g., McDonald’s).\
\&#xNAN;*Benchmark:* McDonald’s rent plus royalty commonly totals **\~12–16%** of sales.

**System Quality Score (QUAL) (aux).** Composite of audits, brand compliance, NPS.\
\&#xNAN;*Pseudo:* `w1*AuditScores + w2*NPS + w3*ComplaintRate`.\
\&#xNAN;*Why it matters:* Brand consistency drives AUV and retention.\
\&#xNAN;*Benchmark:* Franchisors target **>90% pass on audits**; drops trigger remediation.

**Average Unit Volume $ (AUV).** Average annual sales per franchise unit.\
\&#xNAN;*Pseudo:* `Total_Franchisee_Sales / Units`.\
\&#xNAN;*Why it matters:* Drives royalty dollars and franchisee ROI; a key selling point to prospects.\
\&#xNAN;*Benchmark:* McDonald’s U.S. AUV **\~$4 M**; Chick-fil-A **\~$9.3 M**; Subway **\~$0.49 M**.

**Payback Period (Franchisee) (PAYB).** Years for a typical franchisee to recoup initial investment from profits.\
\&#xNAN;*Pseudo:* `Initial_Investment / Annual_Profit`.\
\&#xNAN;*Why it matters:* If franchisees don’t earn back fast enough, growth stalls and churn rises.\
\&#xNAN;*Benchmark:* Many QSRs aim for **<5–7 years**; lower is a strong selling point.




---

[Next Page](/llms-full.txt/1)

