Skip to content

CLI reference

dosi (crate crates/dosi-engine) is the command-line interface to dosi-engine: validate an Apache Ossie (formerly OSI) semantic model, browse its compiled metrics/datasets/dimensions, compile a metric or detail query to dialect SQL, and execute that SQL against a warehouse. Everything the REST API does over HTTP, dosi does from a shell — same compiler, same structured errors, same --format json machine contract.

$ dosi --help
Compile metric queries over OSI semantic models to dialect SQL

Usage: osi [OPTIONS] <COMMAND>

Commands:
  validate  Validate the model: structure, unique names, relationship integrity, and metric compilation
  list      List model objects
  query     Compile a metric query to SQL
  select    Query row-level detail: project fields at one dataset's grain, joining related datasets automatically
  explain   Show the compiled IR shape and logical plan without generating SQL

Installing

dosi ships in two builds: the default carries DuckDB's official libduckdb library beside the executables — files that travel together, nothing else to install — and the lean build drops it, running local DuckDB queries through a duckdb CLI on your PATH instead. Both include all warehouse connectors.

See Install for the choice, the download command, and checksum verification.

Dosi ships for Linux and macOS, and the binaries depend only on standard system libraries (glibc ≥ 2.28 on Linux, macOS 11 or newer); the default build adds libstdc++ through the libduckdb shipped next to it, while the lean build has no C++ dependency at all. TLS in the connectors is rustls, so there is no OpenSSL or database client library to install.

Global options

These apply to every command (clap global = true):

Flag Env Default Meaning
--model <path> DOSI_MODEL required Ossie model file (.yaml / .yml / .json)
--format <fmt> — text Output format: text, json, or arrow (see Output formats)
--connections <path> DOSI_CONNECTIONS discovery chain Connections file for --execute (see Executing against a warehouse)
--osi-datus — on Datus mode: DATUS custom_extensions are honored (datus-extensions.md)
--osi-basic — off Basic mode: strict standard Ossie — DATUS extensions ignored with a warning; validate also runs the upstream validator (see below)

--osi-datus and --osi-basic are mutually exclusive; passing both is a usage error. In basic mode every ignored extension prints a ! warning line on stderr (and lands in warnings under --format json); warnings never change the exit code.

--model may be omitted if DOSI_MODEL is set; a command that needs a model and finds neither exits with no model given; pass --model <path> or set DOSI_MODEL.

Exit codes: 0 success, 1 a query/execute the engine rejected (with a structured error), 2 a usage/CLI error (bad flag, missing model, unparseable argument).

Commands

dosi info

Reports the engine, Ossie spec and datus-ext versions, plus every vendor_name: DATUS custom_extensions key this engine reads. Needs no model, so it doubles as a version probe.

$ dosi info
dosi       0.1.12
osi spec   0.2.0.dev0
datus-ext  1.9 (accepts 1.0 and up)
mode       datus
build      engine (DuckDB in process)
examples   /home/you/.local/share/dosi/examples

KEY               EXTENSION  CARRIERS         SINCE  IF IGNORED
join_type         D-JOIN     relationship     1.0    documented
fill_nulls_with   D-FILL     metric           1.0    documented
time_dimension    D-TIME     dataset, metric  1.1    degraded
time_granularity  D-GRAIN    field            1.1    documented
window            D-WINDOW   metric           1.2    documented
dataset           D-DATASET  metric           1.1    degraded
derive            D-DERIVE   metric           1.4    degraded
measure           D-MEASURE  metric           1.4    silent
params            D-PARAM    metric           1.5    degraded
time              D-FORMAT   field            1.6    silent
is_dimension      D-DIM      field            1.7    documented
cardinality       D-CONFORM  relationship     1.8    silent
lod               D-LOD      field            1.10   degraded

build says which of the two published builds this is — engine runs DuckDB in process, lean shells local DuckDB out to a duckdb CLI on your PATH (see Which download?). Without this line, only ldd can tell an installed binary's two builds apart.

examples is where the bundled example models were found — $DOSI_EXAMPLES if you set it, else $XDG_DATA_HOME/dosi/examples, i.e. ~/.local/share/dosi/examples. The line is absent when neither is there, which is the normal case for a source checkout, where the same models are fixtures/orders and fixtures/tpcds. The examples on this page use $DOSI_EXAMPLES; export it once and every command below is copy-pasteable:

$ export DOSI_EXAMPLES=~/.local/share/dosi/examples

SINCE is the datus-ext version that introduced the key; IF IGNORED is what it costs a consumer that does not honor it — see datus-extensions.md §2.1. --format json emits the same content as a machine contract, which is what a model-generating agent should read before choosing which keys to emit.

dosi validate

Loads the model and runs the full validation pipeline — structure, unique names, relationship reference integrity, and metric-expression compilation — then reports every issue.

$ dosi validate --model $DOSI_EXAMPLES/tpcds/model.yaml
✓ 1 semantic model(s) valid

--format json emits {issues, compile_errors, warnings} for CI gating: each issue carries a severity and message, each compile error a stable code and optional hint, each warning a stable code, location, and message. A model with errors exits non-zero; warnings alone do not.

Field-role advice (datus mode). In --osi-datus — the default — validation also reports fields whose role the model leaves undeclared and structure cannot infer: a measurement no metric aggregates, a business key that is neither a declared key nor a relationship column, a timestamp with no dimension.is_time. Each is a warning that names the field and what to write (usually is_dimension); left alone, those fields are offered as grouping dimensions and an agent will group by them. The check always runs — there is no flag to turn it off — and --strict makes it count, for CI:

$ dosi validate --strict --model model.yaml
! underdeclared_field in field 'routes.distance_miles': `routes.distance_miles` looks like a
  measurement no metric aggregates, so it is offered as a grouping dimension; write the
  metric it belongs to, or declare `is_dimension: false`
✗ the authoring lint reported findings and --strict is set

Upstream validation (--osi-basic). In basic mode validate additionally shells out to the upstream apache/ossie reference validator when a checkout is found — OSSIE_DIR (default ~/src/ossie) containing validation/validate.py — holding the model to the published spec. That validator declares its dependencies as inline script metadata, so it runs through uv run --script when uv is on PATH and through python3 otherwise. An upstream failure fails the command. Anything that stops the validator from running — no checkout, no interpreter, or an interpreter without pyyaml / jsonschema / sqlglot — prints a note and falls back to built-in validation only, because "we could not ask" must not read as "the model is invalid". --format json adds "upstream": {ran, passed, output, note}.

Check the validator out at commit 7b8cdaa, not main: upstream's development schema made a document one unwrapped model in apache/ossie#383, and this release reads only the semantic_model: list form, so upstream main rejects every model it accepts.

$ git clone https://github.com/apache/ossie ~/src/ossie
$ git -C ~/src/ossie checkout 7b8cdaa
$ OSSIE_DIR=~/src/apache-ossie dosi validate --osi-basic --model model.yaml
✓ upstream OSI validation passed
✓ 1 semantic model(s) valid

dosi list <what>

Browse the compiled semantic layer. Three subcommands, each a table in text mode or an array of objects under --format json:

Subcommand Columns
dosi list datasets name, source, primary key, field count, time dimensions
dosi list metrics name, inferred kind (aggregate / ratio / expression), datasets, description
dosi list dimensions [--metric <name>] dataset.field, time flag, whether it is worth grouping by (DIM), who decided (SOURCE), description

list dimensions without --metric lists every field of the model. With --metric it lists only what that metric can be grouped by, recommended first. DIM is no for a field the model declares (or the engine infers) is not a dimension — a measure column, a key, a relationship column — and SOURCE says which: declared for a D-DIM declaration, inferred:<rule> otherwise. A no row is still a legal --group-by; the flag is a recommendation, not a restriction.

$ dosi list dimensions --metric revenue --model $DOSI_EXAMPLES/orders/model.yaml
NAME                   TIME  DIM  SOURCE                DESCRIPTION
metric_time            time  yes  inferred:time
orders.order_date      time  yes  inferred:time
orders.status                yes  inferred
customers.region             yes  inferred
products.category            yes  inferred
orders.amount                no   inferred:measure
orders.customer_id           no   inferred:foreign_key
orders.order_id              no   inferred:primary_key
$ dosi list metrics --model $DOSI_EXAMPLES/tpcds/model.yaml
NAME                     KIND        DATASETS               DESCRIPTION
total_sales              aggregate   store_sales            Total sales revenue across all transactions
total_profit             aggregate   store_sales            Total net profit from store sales
customer_lifetime_value  ratio       customer, store_sales  Average lifetime sales value per customer
sales_by_brand           aggregate   store_sales            Total sales by brand (requires grouping by item.i_brand)
store_productivity       expression  store, store_sales     Sales per employee across stores

dosi query

The core command: compile a metric query into dialect SQL, and optionally run it. The query is described by the shared query spec flags below; on top of those, query takes:

Flag Meaning
--dialect <name> Target SQL dialect (default duckdb, or the --connection profile's dialect)
--pretty Pretty-print the generated SQL
--explain Also print the logical plan above the SQL
--execute Run the compiled SQL against a warehouse and print the rows
--connection <name> Named profile from the connections file (implies its dialect)
--db <path> DuckDB file for --execute without --connection (default: in-memory)

Compile only (no --execute) prints the SQL (text) or {dialect, sql} (json):

$ dosi query --model $DOSI_EXAMPLES/tpcds/model.yaml \
    --metrics total_sales \
    --group-by store.s_state,date_dim.d_date:month \
    --where "item.i_category = 'Books'" \
    --start-time 2024-01-01 --end-time 2025-01-01 \
    --order -total_sales --limit 100 \
    --dialect starrocks
SELECT store.s_state AS s_state, DATE_TRUNC('MONTH', date_dim.d_date) AS d_date__month, ...

(Combining total_sales with a store-based metric like store_productivity here would be a fan_out_risk error — a per-store measure can't be grouped by a date the store doesn't reach. That protection is the point; see semantics.md §6.)

Query spec flags (shared with explain)

Flag Meaning
--metrics <a,b,…> Required. Comma-separated metric names
--group-by <items> Comma-separated dataset.field or dataset.field:grain (grain: day\|week\|month\|quarter\|year); metric_time[:grain] groups by each metric's primary time dimension
--where <sql> Scalar boolean SQL, placed per AND conjunct: a condition on a selected metric filters the result rows (before --order / --limit); in a window query so does one on group-by fields only, so ranks and partition counts stay those of the whole population; anything else filters rows before aggregation. See filters.
--context-filter <sql> Row-level SQL that always filters before aggregation — the population a window ranks and counts. Use it to rank within a subset.
--start-time <YYYY-MM-DD> Inclusive lower bound of a time range
--end-time <YYYY-MM-DD> Exclusive upper bound
--time-dimension <field> Which time dimension the range applies to (default: the only time dimension in the group-by, else each metric's primary time; metric_time says so explicitly)
--order <keys> Comma-separated order keys; a - prefix means descending
--limit <n> Row limit
--param <name=value[,value…]> Bind a D-PARAM parameter (repeatable). Values are typed by the metric's declaration; a comma list expands the metric into one column per value; quote a string value that contains a comma. See datus-extensions.md

Three behaviors worth internalizing (the full contract is in semantics.md):

  • Grain output naming. --group-by orders.order_date:month produces a column named order_date__month ({field}__{grain}). This is also what you reference in --order.
  • Order keys are output column names, not qualified fields. --order -total_sales or --order ds__month — never orders.status. The - prefix descends; because clap would read it as a flag, --order allows leading hyphens.
  • Time ranges are half-open [start, end). --start-time 2024-01-01 --end-time 2025-01-01 includes all of 2024 and excludes 2025-01-01 exactly. With no --time-dimension and no time item in the group-by, the range falls back to each metric's primary time dimension (declared via the Datus D-TIME extension, or the dataset's single is_time field — see datus-extensions.md); only a group-by holding several time dimensions still needs an explicit --time-dimension (time_range_needs_dimension). The reserved name metric_time selects the primary time explicitly, in --group-by (metric_time:month → output column metric_time__month) and --time-dimension alike.

dosi select

Compiles a detail query: row-level records at one dataset's grain, rather than aggregated metrics. Same planner, same relationship graph, same structured errors — see Detail queries for the full guide.

$ dosi select --model $DOSI_EXAMPLES/orders/model.yaml \
    --from orders \
    --fields orders.order_id,orders.amount,customers.name \
    --where "customers.tier = 'Enterprise'" \
    --order -amount --limit 20 --execute

It takes the same --dialect / --pretty / --explain / --execute / --connection / --db flags as query, plus its own spec:

Flag Meaning
--from <dataset> Required. The result grain: one row per row of this dataset. Not a SQL FROM clause — other datasets are joined automatically.
--fields <a,b,...> Required. field, dataset.field, or relationship[.relationship…].field, each with an optional :grain on a time field.
--where <sql> Row-level filter. May reference datasets the projection does not, which joins them. Aggregates are rejected.
--start-time / --end-time Half-open [start, end) window.
--time-dimension <field> Which time field the window applies to; defaults to the root dataset's primary time dimension.
--order <k,-k,...> Order keys, naming projected columns (- prefix = descending).
--limit <n> Default 100, clamped to 10000.

Three behaviors worth internalizing:

  • Output columns drop the root prefix and join the rest with __: orders.amount is amount, customers.name is customers__name.
  • A one-to-many traversal is refused, not silently duplicated. Asking --from customers --fields orders.order_id returns detail_fanout with the root to use instead.
  • When two relationship paths reach a field, name the one you mean by prefixing it — the ambiguous_join_path error's suggested_retry prints both spellings ready to paste. The same prefix works in query --group-by.

dosi explain

Takes the same query spec flags as query but stops at the logical plan — no SQL is generated. Useful for understanding join paths, fan-out branch assignment, and grain handling before you pick a dialect.

$ dosi explain --model $DOSI_EXAMPLES/tpcds/model.yaml \
    --metrics customer_lifetime_value --group-by store.s_state

The plan renders as text (including each join's kind, left / inner, so you can confirm a Datus join_type extension took effect). --format json applies only to the error path here — a successful plan is text-only.

dosi attribute

Decomposes a metric's change between two [start, end) windows into per-dimension contributions — the engine runs every needed query itself and picks the method exact for the metric's type. See Attribution for what comes back and how to read it.

$ dosi attribute --model model.yaml \
    --metric revenue --dimensions status,customers.region \
    --baseline 2024-01-01..2024-02-01 --current 2024-02-01..2024-03-01 \
    --db warehouse.duckdb

Windows are START..END ISO ranges. --where scopes the whole analysis; --connection/--db pick the warehouse exactly like query --execute; --max-values, --top-dimensions, and --top-values bound the output. The result is always JSON.

dosi lineage

The metric lineage graph — physical tables → datasets → atomic metrics → derived metrics, and the ontology's concepts when one is loaded — as JSON or as a page you can open. See Lineage for what the graph contains.

$ dosi lineage dump --model model.yaml > graph.json
$ dosi lineage view --model model.yaml --ontology ontology.yaml

dump prints the graph JSON (--out <path> writes it to a file; the contract is JSON, so --format does not apply). view writes a single self-contained HTML page and opens it in your default browser; --out <path> picks where the page goes, and --no-open — or a machine with no browser — prints the path instead and still exits 0. The page needs no network access and opens over file://.

Both take --redact-sql, which replaces every SQL text and opaque extension payload with "<redacted>" while keeping the shape, for sharing a model's structure without its expressions. --osi-basic compiles under strict OSI and stamps "mode": "basic" on the output.

Output formats

--format takes one of three values:

  • text (default) — human-readable: an aligned table for rows/listings, the raw SQL for a compile, an indented tree for a plan. NULL cells render dimmed. Colors auto-detect the terminal.
  • json — every output (and every error) is machine-readable. A query compile is {dialect, sql}; --execute adds {columns, rows: [{col: val}], …}; validate is {issues, compile_errors}; list is an array of objects. Errors carry a stable code, the names involved, candidates where a bad reference has alternatives, and a suggested_retry where a rewrite would succeed — so agentic callers self-correct without parsing prose. Error codes are the stable API; error text is not.
  • arrow — an Arrow IPC stream on stdout, for query --execute only (any other command errors: --format arrow only applies to 'query --execute'). Result batches pass straight from the warehouse adapter to stdout with no row materialization — pipe them into DuckDB, Polars, or pyarrow with zero JSON parsing. Requires an Arrow-capable build (the default build, or any exec-*-arrow / exec-flightsql / exec-duckdb feature; a --no-default-features build without one rejects --format arrow at runtime).
# Stream results into DuckDB for further analysis
dosi query --model model.yaml --metrics revenue --group-by orders.status \
    --execute --connection prod-ch --format arrow \
    | duckdb -c "SELECT * FROM read_arrow('/dev/stdin')"

# Or into Polars
dosi query ... --execute --format arrow \
    | python -c "import polars as pl,sys; print(pl.read_ipc_stream(sys.stdin.buffer))"

Executing against a warehouse

--execute runs the compiled SQL and prints the rows. Without a connection it uses local DuckDB — in-process and Arrow-native by default (bundled exec-duckdb); --db <file> points at a DuckDB file, otherwise it is in-memory. With --connection <name> it targets that profile's warehouse and dialect.

$ dosi query --model model.yaml --metrics revenue --group-by orders.status \
    --execute --connection prod-sr

Connection profiles use the Datus agent.yml datasources: vocabulary. Point --connections at a full agent.yml (reads services.datasources) or a standalone datasources: file. Without the flag, the file is discovered in order: DOSI_CONNECTIONS env → ./dosi-connections.yaml → ./osi-connections.yaml → ~/.config/dosi/connections.yaml → ~/.config/osi/connections.yaml → ./conf/agent.yml → ~/.datus/conf/agent.yml — an existing Datus install works with zero config, and the pre-rename osi- paths still resolve. Secrets interpolate from the environment as ${VAR}.

datasources:
  prod-sr:
    type: starrocks
    host: sr.internal
    port: 9030
    arrow_flight_port: 9408      # opt into Arrow Flight SQL (SR ≥3.5.1)
    username: osi
    password: ${SR_PASSWORD}
    database: analytics
    default: true                # used by --execute without --connection
  prod-ch:
    type: clickhouse
    uri: http://ch.internal:8123
    username: default
    database: analytics

Resolution rules:

  • --execute without --connection uses the profile marked default: true; if none is marked, it falls back to local DuckDB (--db or in-memory). A default: true entry that failed to parse (e.g. an unset ${VAR}) warns on stderr rather than silently using empty DuckDB.
  • --dialect with --connection must agree with the profile's dialect, or the command errors — drop --dialect and let the profile decide.

Warehouse drivers are feature-gated if you build from source, so a hand-rolled binary only carries what you ask for (published builds carry them all). Build with the features you need (or exec-all):

Feature Engines Result path
exec-duckdb (default) DuckDB (in-process) Arrow-native
exec-mysql MySQL, TiDB, StarRocks, Doris (MySQL wire) rows
exec-postgres Postgres rows
exec-hologres Hologres (Postgres wire; implies exec-postgres) rows
exec-gaussdb GaussDB / openGauss (native SHA256 auth driver) rows
exec-oracle Oracle Database (ODPI-C; Instant Client at runtime) rows
exec-http ClickHouse, Trino rows
exec-http-arrow ClickHouse FORMAT ArrowStream Arrow-native
exec-flightsql StarRocks / Doris Arrow Flight SQL (arrow_flight_port:) Arrow-native
exec-snowflake Snowflake (SQL API v2 + key-pair JWT) rows
exec-bigquery BigQuery (jobs.query REST + service-account JWT) rows

Per-engine setup and the Arrow result-path status live in connectors.md and arrow.md.

See also

  • semantics.md — the normative behavior contract (metric inference, joins, fan-out protection, time handling) the CLI compiles to.
  • rest-api.md — the same capabilities over HTTP.
  • extensions-guide.md — the optional Datus model extensions (join_type, fill_nulls_with, time_dimension, time_granularity, time, dataset), and datus-extensions.md for their normative contract and versioning.
  • connectors.md — warehouse connector setup.
  • arrow.md — enabling Arrow result transfer and what gets faster.