Skip to content

What Is Apache Ossie? And Why Dosi?

Apache Ossie, formerly known as Open Semantic Interchange (OSI), is an open specification for exchanging semantic models, metrics, dimensions, and business context across BI tools, AI agents, and data platforms.

An Ossie model is a plain YAML file. Define a metric in it once, and every tool that speaks Ossie reads that one definition instead of re-implementing it. Dosi is the first Apache Ossie-native semantic layer engine: it loads an Ossie model, validates it against the upstream spec, and compiles metric requests into correct SQL for 16 warehouse dialects.

Apache Ossie at a glance

What it is An open, vendor-neutral specification for semantic models
Formerly Open Semantic Interchange (OSI)
Format YAML (or JSON)
Describes Datasets, relationships, fields and dimensions, metrics, and AI context for each
Leaves to engines How to turn a metric into correct SQL on a specific warehouse
Upstream ossie.apache.org and apache/ossie: the core spec, a reference validator, and format converters
Engine on this site Dosi: validates, compiles, and runs Ossie models

The problem: everyone redefines "revenue"

Write "August revenue" as SQL from scratch and there are three separate ways to get it wrong: the join (which table has revenue — and does joining it fan out the rows and quietly double the total?), the dimension (order date or ship date, gross or net, daily grain or monthly), and the expression (SUM(amount), or SUM(unit_price * quantity) from a line-items table that skips the discount column, or SUM(CASE WHEN status != 'refunded' THEN amount ELSE 0 END) to back out refunds). Miss any one and the result still looks plausible.

That's exactly why a BI dashboard, a finance spreadsheet, and a hand-written query end up as three different "revenues" — one excludes cancelled orders, another doesn't, a third's join doubles the total nobody checked. And it's why an AI agent asked to write the query from raw table names makes the same three mistakes, just with more confidence. The result is familiar to everyone: two "correct" numbers that don't match, and a meeting to figure out which one to trust.

A semantic model fixes this by defining each metric once — what revenue means, which table it comes from, how tables relate — as a single source of truth. Every query then derives from that one definition instead of re-implementing it. Apache Ossie is an open, standard way to write that model down.

What an Apache Ossie model looks like

A model lists your datasets (tables) and their fields, the relationships between them, and your metrics as ordinary SQL expressions. This one is complete and valid:

version: "0.2.0.dev0"

semantic_model:
  - name: orders_model
    description: Orders, the customers who placed them, and revenue
    datasets:
      - name: orders
        source: main.orders
        primary_key: [order_id]
        fields:
          - name: order_id
            expression:
              dialects:
                - dialect: ANSI_SQL
                  expression: order_id
          - name: customer_id
            expression:
              dialects:
                - dialect: ANSI_SQL
                  expression: customer_id
          - name: order_date
            expression:
              dialects:
                - dialect: ANSI_SQL
                  expression: order_date
            dimension:
              is_time: true
          - name: amount
            expression:
              dialects:
                - dialect: ANSI_SQL
                  expression: amount
      - name: customers
        source: main.customers
        primary_key: [customer_id]
        fields:
          - name: customer_id
            expression:
              dialects:
                - dialect: ANSI_SQL
                  expression: customer_id
          - name: region
            expression:
              dialects:
                - dialect: ANSI_SQL
                  expression: region

    relationships:
      - name: orders_to_customers
        from: orders
        to: customers
        from_columns: [customer_id]
        to_columns: [customer_id]

    metrics:
      - name: revenue
        description: Total order amount
        expression:
          dialects:
            - dialect: ANSI_SQL
              expression: SUM(orders.amount)
        ai_context:
          synonyms: [sales, order revenue]
  • Datasets and fields name the physical tables and the columns you can group and filter by; dimension: {is_time: true} marks a time dimension.
  • Relationships say how tables join, so a consumer never has to guess the join key.
  • Metrics are SQL aggregate expressions over those fields, written once.
  • ai_context carries business context — synonyms, instructions, examples — for AI agents that read the model.

Ossie describes what your metrics are. It deliberately does not say how to turn them into correct SQL for a specific warehouse. That is the gap Dosi fills.

An interchange, not just a format

The Interchange in the original name still matters: the upstream Apache Ossie project maintains bidirectional converters between Ossie and other semantic-layer formats — dbt, Snowflake, Databricks, Salesforce/Tableau, GoodData, Honeydew, Omni, and more. One neutral hub instead of point-to-point migrations.

Dosi reads standard Ossie models as they are, so a converter's output works unchanged as long as its metrics are SQL expressions (see S-EXPR-1). A model can arrive from dbt today and leave for Snowflake tomorrow; your metric definitions are never hostage to one vendor's format. Define once, and move freely.

How Dosi relates to Apache Ossie

Apache Ossie is the contract; Dosi is the engine that executes it. Dosi does not define its own model format: its input is an Ossie model, and everything it adds either happens at query time or rides inside the extension point the spec itself provides.

Apache Ossie defines Dosi adds
Datasets, relationships, fields, and metrics as SQL expressions Join planning with fan-out protection, and SQL for 16 dialects
A reference validator dosi validate --osi-basic runs it alongside Dosi's own checks
custom_extensions, the spec's vendor extension point The Datus extensions — windows, derived metrics, join type, null-fill, time grain — carried inside it
AI context on every object A native MCP server that serves the model to agents, with structured errors
Converters to and from other formats Converted models with SQL metrics accepted as input, unchanged

Dosi runs in one of two modes, chosen per invocation:

  • Basic (--osi-basic): strict standard Apache Ossie. The model means exactly what the spec says; Datus extensions are ignored with a warning.
  • Datus (--osi-datus, the default): Apache Ossie plus the Datus extensions.

Datus extensions live inside custom_extensions, so they never make a model invalid: other Ossie tools read it and skip the extensions they don't know. Datus mode is more lenient than the spec in one respect — it also accepts vendor SQL dialect tags such as POSTGRESQL that the core spec does not define — so run dosi validate --osi-basic when a model has to stay valid for other Ossie tools.

Dosi is developed by Datus. It implements the Apache Ossie specification but is not itself an Apache Software Foundation project.

What Dosi does

Dosi is a small, fast (Rust) engine that reads a pure Apache Ossie model and turns a metric request into correct, warehouse-specific SQL — then, optionally, runs it:

flowchart LR
    A["Your Apache Ossie model<br/><small>metrics defined once</small>"] --> B["Dosi"]
    B --> C["Correct SQL<br/><small>for your warehouse</small>"]
    C --> D["Results"]

You ask for a metric, a few dimensions to group by, and a warehouse dialect. Dosi works out the joins, the aggregation, and the exact SQL idioms that dialect needs — the same model produces correct SQL for 16 dialects: DuckDB, Postgres, MySQL, ClickHouse, Snowflake, BigQuery, StarRocks, Trino, Databricks, Oracle, and more. Switching warehouses changes one flag, not your metric definitions.

You can use it four ways — a command-line tool, a REST + Arrow server, a native MCP server for AI agents, or Python bindings — all sharing the same engine and the same answers.

What an AI agent needs from a metric system

When the consumer of a metric is an AI agent rather than a person reading a dashboard, the metric layer has to answer questions a human analyst normally settles by judgment. For every metric, four properties decide what can be done with it safely:

  • Additive — can it be summed across a dimension, or would adding it double-count (distinct counts, ratios)?
  • Rollable — do daily values roll up into a monthly one, or must the month be recomputed from raw rows?
  • Predictable — is it a flow you can forecast forward, or a stock where extrapolation is meaningless?
  • Fan-out-safe — will a join multiply rows and silently inflate the total?

Dosi treats these as properties the engine must know, not conventions in an analyst's head. Today it infers every metric's kind (aggregate / ratio / expression) and never silently double-counts: when a metric would join tables at different levels of detail — the classic fan-out that doubles your revenue — it either computes each part at its own grain, or stops with a structured error. It will not hand back a quietly inflated total.

Built-in attribution, chosen per metric kind: when an agent asks "why did revenue drop?", one call computes the contribution breakdown with the algorithm that is exact for that metric's algebra — dimension attribution for additive metrics, an LMDI mix/rate decomposition for ratios — so the agent's reasoning is observable and checkable, not hallucinated.

On the authoring side, datus-agent — Datus' own open-source data agent — is the best first-party way to produce these models: it generates Dosi-ready Apache Ossie YAML with the Datus extensions, which add more advanced window and derived metrics and more complete join-relationship detection than the base spec. The consumption side stays open: any agent can query the result through the MCP server — Claude Code, for example.

What makes the numbers trustworthy

Beyond the metric algebra above, the engine closes off the everyday mistakes that produce mismatched numbers:

  • Time ranges are unambiguous. A range like Jan 2024 means [2024-01-01, 2025-01-01) — start included, end excluded — so days never get double-counted at month boundaries.
  • Errors tell you how to fix them. A wrong metric name doesn't get a stack trace; it gets a message naming the closest valid options. In --format json, every error carries a stable code and suggested fix, so automated tools (and AI agents) can self-correct.

These guarantees are written down precisely — as a normative contract — in the semantics reference. You don't need to read it to use Dosi, but it's there when you want to know exactly what the engine promises.

Who it's for

  • Analytics & data engineers who want one metric definition that stays correct across every warehouse and every consumer.
  • Teams migrating or multi-homing warehouses who don't want to rewrite metric SQL per engine.
  • Teams standardizing on Apache Ossie who need an engine that runs their models as they are, without a proprietary format in between.
  • Builders of data apps and AI agents who need a metric API with structured, machine-readable errors instead of free-form SQL.

FAQ

Is OSI the same as Apache Ossie?

Yes. Open Semantic Interchange (OSI) is the specification's former name; Apache Ossie is its current one. Models, tooling, and documentation that say OSI refer to the same spec.

Does Dosi work with any Apache Ossie model?

Any model whose metrics are written in SQL. Ossie also allows MDX, TABLEAU and MAQL expressions; Dosi compiles SQL only, so a metric with no SQL entry is a compile error and a field with none is a warning (see S-EXPR-1). dosi validate --osi-basic checks a model with the upstream Apache Ossie validator as well as Dosi's own rules. The Datus extensions are optional.

Why do some Dosi flags still say osi?

--osi-basic, --osi-datus, and a few environment variables were named before the rename. They keep their names so existing scripts and configurations keep working.

Next steps