CLI reference¶
dosi (crate crates/dosi-engine) is the command-line
interface to dosi-engine: validate an
Apache Ossie (formerly OSI) semantic model,
browse its compiled metrics/datasets/dimensions, compile a metric or
detail query to dialect SQL, and execute that SQL against a
warehouse. Everything the
REST API does over HTTP, dosi does from a shell — same
compiler, same structured errors, same --format json machine contract.
$ dosi --help
Compile metric queries over OSI semantic models to dialect SQL
Usage: osi [OPTIONS] <COMMAND>
Commands:
validate Validate the model: structure, unique names, relationship integrity, and metric compilation
list List model objects
query Compile a metric query to SQL
select Query row-level detail: project fields at one dataset's grain, joining related datasets automatically
explain Show the compiled IR shape and logical plan without generating SQL
Installing¶
dosi ships in two builds: the default carries DuckDB's official libduckdb
library beside the executables — files that travel together, nothing else to
install — and the lean build drops it, running local DuckDB queries through
a duckdb CLI on your PATH instead. Both include all warehouse connectors.
See Install for the choice, the download command, and checksum verification.
Dosi ships for Linux and macOS, and the binaries depend only on standard system
libraries (glibc ≥ 2.28 on Linux, macOS 11 or newer); the default build adds
libstdc++ through the libduckdb shipped next to it, while the lean build has
no C++ dependency at all. TLS in the
connectors is rustls, so there is no OpenSSL or database client library to
install.
Global options¶
These apply to every command (clap global = true):
| Flag | Env | Default | Meaning |
|---|---|---|---|
--model <path> |
DOSI_MODEL |
required | Ossie model file (.yaml / .yml / .json) |
--format <fmt> |
— | text |
Output format: text, json, or arrow (see Output formats) |
--connections <path> |
DOSI_CONNECTIONS |
discovery chain | Connections file for --execute (see Executing against a warehouse) |
--osi-datus |
— | on | Datus mode: DATUS custom_extensions are honored (datus-extensions.md) |
--osi-basic |
— | off | Basic mode: strict standard Ossie — DATUS extensions ignored with a warning; validate also runs the upstream validator (see below) |
--osi-datus and --osi-basic are mutually exclusive; passing both is a
usage error. In basic mode every ignored extension prints a ! warning line
on stderr (and lands in warnings under --format json); warnings never
change the exit code.
--model may be omitted if DOSI_MODEL is set; a command that needs a model
and finds neither exits with no model given; pass --model <path> or set
DOSI_MODEL.
Exit codes: 0 success, 1 a query/execute the engine rejected (with a
structured error), 2 a usage/CLI error (bad flag, missing model, unparseable
argument).
Commands¶
dosi info¶
Reports the engine, Ossie spec and datus-ext versions, plus every
vendor_name: DATUS custom_extensions key this engine reads. Needs no model,
so it doubles as a version probe.
$ dosi info
dosi 0.1.12
osi spec 0.2.0.dev0
datus-ext 1.9 (accepts 1.0 and up)
mode datus
build engine (DuckDB in process)
examples /home/you/.local/share/dosi/examples
KEY EXTENSION CARRIERS SINCE IF IGNORED
join_type D-JOIN relationship 1.0 documented
fill_nulls_with D-FILL metric 1.0 documented
time_dimension D-TIME dataset, metric 1.1 degraded
time_granularity D-GRAIN field 1.1 documented
window D-WINDOW metric 1.2 documented
dataset D-DATASET metric 1.1 degraded
derive D-DERIVE metric 1.4 degraded
measure D-MEASURE metric 1.4 silent
params D-PARAM metric 1.5 degraded
time D-FORMAT field 1.6 silent
is_dimension D-DIM field 1.7 documented
cardinality D-CONFORM relationship 1.8 silent
lod D-LOD field 1.10 degraded
build says which of the two published builds this is — engine runs DuckDB
in process, lean shells local DuckDB out to a duckdb CLI on your PATH
(see Which download?). Without this line, only
ldd can tell an installed binary's two builds apart.
examples is where the bundled example models were found — $DOSI_EXAMPLES
if you set it, else $XDG_DATA_HOME/dosi/examples, i.e.
~/.local/share/dosi/examples. The line is absent when neither is there, which
is the normal case for a source checkout, where the same models are
fixtures/orders and fixtures/tpcds. The examples on this page use $DOSI_EXAMPLES; export it
once and every command below is copy-pasteable:
SINCE is the datus-ext version that introduced the key; IF IGNORED is what
it costs a consumer that does not honor it — see
datus-extensions.md §2.1.
--format json emits the same content as a machine contract, which is what a
model-generating agent should read before choosing which keys to emit.
dosi validate¶
Loads the model and runs the full validation pipeline — structure, unique names, relationship reference integrity, and metric-expression compilation — then reports every issue.
--format json emits {issues, compile_errors, warnings} for CI gating: each
issue carries a severity and message, each compile error a stable code and
optional hint, each warning a stable code, location, and message. A
model with errors exits non-zero; warnings alone do not.
Field-role advice (datus mode). In --osi-datus — the default — validation
also reports fields whose role the model leaves undeclared and structure cannot
infer: a measurement no metric aggregates, a business key that is neither a
declared key nor a relationship column, a timestamp with no
dimension.is_time. Each is a warning that names the field and what to write
(usually is_dimension); left alone, those fields
are offered as grouping dimensions and an agent will group by them. The check
always runs — there is no flag to turn it off — and --strict makes it count,
for CI:
$ dosi validate --strict --model model.yaml
! underdeclared_field in field 'routes.distance_miles': `routes.distance_miles` looks like a
measurement no metric aggregates, so it is offered as a grouping dimension; write the
metric it belongs to, or declare `is_dimension: false`
✗ the authoring lint reported findings and --strict is set
Upstream validation (--osi-basic). In basic mode validate additionally
shells out to the upstream apache/ossie
reference validator when a checkout is found — OSSIE_DIR (default
~/src/ossie) containing validation/validate.py — holding the model to the
published spec. That validator declares its dependencies as inline script
metadata, so it runs through uv run --script when uv
is on PATH and through python3 otherwise. An upstream failure fails the
command. Anything that stops the validator from running — no checkout, no
interpreter, or an interpreter without pyyaml / jsonschema / sqlglot —
prints a note and falls back to built-in validation only, because "we could
not ask" must not read as "the model is invalid". --format json adds
"upstream": {ran, passed, output, note}.
Check the validator out at commit 7b8cdaa, not main: upstream's
development schema made a document one unwrapped model in apache/ossie#383,
and this release reads only the semantic_model: list form, so upstream
main rejects every model it accepts.
$ OSSIE_DIR=~/src/apache-ossie dosi validate --osi-basic --model model.yaml
✓ upstream OSI validation passed
✓ 1 semantic model(s) valid
dosi list <what>¶
Browse the compiled semantic layer. Three subcommands, each a table in
text mode or an array of objects under --format json:
| Subcommand | Columns |
|---|---|
dosi list datasets |
name, source, primary key, field count, time dimensions |
dosi list metrics |
name, inferred kind (aggregate / ratio / expression), datasets, description |
dosi list dimensions [--metric <name>] |
dataset.field, time flag, whether it is worth grouping by (DIM), who decided (SOURCE), description |
list dimensions without --metric lists every field of the model. With
--metric it lists only what that metric can be grouped by, recommended first.
DIM is no for a field the model declares (or the engine infers) is not a
dimension — a measure column, a key, a relationship column — and SOURCE says
which: declared for a D-DIM declaration,
inferred:<rule> otherwise. A no row is still a legal --group-by; the
flag is a recommendation, not a restriction.
$ dosi list dimensions --metric revenue --model $DOSI_EXAMPLES/orders/model.yaml
NAME TIME DIM SOURCE DESCRIPTION
metric_time time yes inferred:time
orders.order_date time yes inferred:time
orders.status yes inferred
customers.region yes inferred
products.category yes inferred
orders.amount no inferred:measure
orders.customer_id no inferred:foreign_key
orders.order_id no inferred:primary_key
$ dosi list metrics --model $DOSI_EXAMPLES/tpcds/model.yaml
NAME KIND DATASETS DESCRIPTION
total_sales aggregate store_sales Total sales revenue across all transactions
total_profit aggregate store_sales Total net profit from store sales
customer_lifetime_value ratio customer, store_sales Average lifetime sales value per customer
sales_by_brand aggregate store_sales Total sales by brand (requires grouping by item.i_brand)
store_productivity expression store, store_sales Sales per employee across stores
dosi query¶
The core command: compile a metric query into dialect SQL, and optionally run
it. The query is described by the shared query spec flags below; on top of
those, query takes:
| Flag | Meaning |
|---|---|
--dialect <name> |
Target SQL dialect (default duckdb, or the --connection profile's dialect) |
--pretty |
Pretty-print the generated SQL |
--explain |
Also print the logical plan above the SQL |
--execute |
Run the compiled SQL against a warehouse and print the rows |
--connection <name> |
Named profile from the connections file (implies its dialect) |
--db <path> |
DuckDB file for --execute without --connection (default: in-memory) |
Compile only (no --execute) prints the SQL (text) or {dialect, sql}
(json):
$ dosi query --model $DOSI_EXAMPLES/tpcds/model.yaml \
--metrics total_sales \
--group-by store.s_state,date_dim.d_date:month \
--where "item.i_category = 'Books'" \
--start-time 2024-01-01 --end-time 2025-01-01 \
--order -total_sales --limit 100 \
--dialect starrocks
SELECT store.s_state AS s_state, DATE_TRUNC('MONTH', date_dim.d_date) AS d_date__month, ...
(Combining total_sales with a store-based metric like store_productivity
here would be a fan_out_risk error — a per-store measure can't be grouped by
a date the store doesn't reach. That protection is the point; see
semantics.md §6.)
Query spec flags (shared with explain)¶
| Flag | Meaning |
|---|---|
--metrics <a,b,…> |
Required. Comma-separated metric names |
--group-by <items> |
Comma-separated dataset.field or dataset.field:grain (grain: day\|week\|month\|quarter\|year); metric_time[:grain] groups by each metric's primary time dimension |
--where <sql> |
Scalar boolean SQL, placed per AND conjunct: a condition on a selected metric filters the result rows (before --order / --limit); in a window query so does one on group-by fields only, so ranks and partition counts stay those of the whole population; anything else filters rows before aggregation. See filters. |
--context-filter <sql> |
Row-level SQL that always filters before aggregation — the population a window ranks and counts. Use it to rank within a subset. |
--start-time <YYYY-MM-DD> |
Inclusive lower bound of a time range |
--end-time <YYYY-MM-DD> |
Exclusive upper bound |
--time-dimension <field> |
Which time dimension the range applies to (default: the only time dimension in the group-by, else each metric's primary time; metric_time says so explicitly) |
--order <keys> |
Comma-separated order keys; a - prefix means descending |
--limit <n> |
Row limit |
--param <name=value[,value…]> |
Bind a D-PARAM parameter (repeatable). Values are typed by the metric's declaration; a comma list expands the metric into one column per value; quote a string value that contains a comma. See datus-extensions.md |
Three behaviors worth internalizing (the full contract is in semantics.md):
- Grain output naming.
--group-by orders.order_date:monthproduces a column namedorder_date__month({field}__{grain}). This is also what you reference in--order. - Order keys are output column names, not qualified fields.
--order -total_salesor--order ds__month— neverorders.status. The-prefix descends; because clap would read it as a flag,--orderallows leading hyphens. - Time ranges are half-open
[start, end).--start-time 2024-01-01 --end-time 2025-01-01includes all of 2024 and excludes 2025-01-01 exactly. With no--time-dimensionand no time item in the group-by, the range falls back to each metric's primary time dimension (declared via the Datus D-TIME extension, or the dataset's singleis_timefield — see datus-extensions.md); only a group-by holding several time dimensions still needs an explicit--time-dimension(time_range_needs_dimension). The reserved namemetric_timeselects the primary time explicitly, in--group-by(metric_time:month→ output columnmetric_time__month) and--time-dimensionalike.
dosi select¶
Compiles a detail query: row-level records at one dataset's grain, rather than aggregated metrics. Same planner, same relationship graph, same structured errors — see Detail queries for the full guide.
$ dosi select --model $DOSI_EXAMPLES/orders/model.yaml \
--from orders \
--fields orders.order_id,orders.amount,customers.name \
--where "customers.tier = 'Enterprise'" \
--order -amount --limit 20 --execute
It takes the same --dialect / --pretty / --explain /
--execute / --connection / --db flags as query, plus its own spec:
| Flag | Meaning |
|---|---|
--from <dataset> |
Required. The result grain: one row per row of this dataset. Not a SQL FROM clause — other datasets are joined automatically. |
--fields <a,b,...> |
Required. field, dataset.field, or relationship[.relationship…].field, each with an optional :grain on a time field. |
--where <sql> |
Row-level filter. May reference datasets the projection does not, which joins them. Aggregates are rejected. |
--start-time / --end-time |
Half-open [start, end) window. |
--time-dimension <field> |
Which time field the window applies to; defaults to the root dataset's primary time dimension. |
--order <k,-k,...> |
Order keys, naming projected columns (- prefix = descending). |
--limit <n> |
Default 100, clamped to 10000. |
Three behaviors worth internalizing:
- Output columns drop the root prefix and join the rest with
__:orders.amountisamount,customers.nameiscustomers__name. - A one-to-many traversal is refused, not silently duplicated. Asking
--from customers --fields orders.order_idreturnsdetail_fanoutwith the root to use instead. - When two relationship paths reach a field, name the one you mean by
prefixing it — the
ambiguous_join_patherror'ssuggested_retryprints both spellings ready to paste. The same prefix works inquery --group-by.
dosi explain¶
Takes the same query spec flags as query but stops at the logical plan — no
SQL is generated. Useful for understanding join paths, fan-out branch
assignment, and grain handling before you pick a dialect.
$ dosi explain --model $DOSI_EXAMPLES/tpcds/model.yaml \
--metrics customer_lifetime_value --group-by store.s_state
The plan renders as text (including each join's kind, left / inner, so you
can confirm a Datus join_type extension took effect).
--format json applies only to the error path here — a successful plan is
text-only.
dosi attribute¶
Decomposes a metric's change between two [start, end) windows into
per-dimension contributions — the engine runs every needed query itself and
picks the method exact for the metric's type. See
Attribution for what comes back and how to read it.
$ dosi attribute --model model.yaml \
--metric revenue --dimensions status,customers.region \
--baseline 2024-01-01..2024-02-01 --current 2024-02-01..2024-03-01 \
--db warehouse.duckdb
Windows are START..END ISO ranges. --where scopes the whole analysis;
--connection/--db pick the warehouse exactly like query --execute;
--max-values, --top-dimensions, and --top-values bound the output.
The result is always JSON.
dosi lineage¶
The metric lineage graph — physical tables → datasets → atomic metrics → derived metrics, and the ontology's concepts when one is loaded — as JSON or as a page you can open. See Lineage for what the graph contains.
$ dosi lineage dump --model model.yaml > graph.json
$ dosi lineage view --model model.yaml --ontology ontology.yaml
dump prints the graph JSON (--out <path> writes it to a file; the
contract is JSON, so --format does not apply). view writes a
single self-contained HTML page and opens it in your default browser;
--out <path> picks where the page goes, and --no-open — or a machine with
no browser — prints the path instead and still exits 0. The page needs no
network access and opens over file://.
Both take --redact-sql, which replaces every SQL text and opaque
extension payload with "<redacted>" while keeping the shape, for sharing a
model's structure without its expressions. --osi-basic compiles under
strict OSI and stamps "mode": "basic" on the output.
Output formats¶
--format takes one of three values:
text(default) — human-readable: an aligned table for rows/listings, the raw SQL for a compile, an indented tree for a plan.NULLcells render dimmed. Colors auto-detect the terminal.json— every output (and every error) is machine-readable. Aquerycompile is{dialect, sql};--executeadds{columns, rows: [{col: val}], …};validateis{issues, compile_errors};listis an array of objects. Errors carry a stablecode, the names involved,candidateswhere a bad reference has alternatives, and asuggested_retrywhere a rewrite would succeed — so agentic callers self-correct without parsing prose. Error codes are the stable API; error text is not.arrow— an Arrow IPC stream on stdout, forquery --executeonly (any other command errors:--format arrow only applies to 'query --execute'). Result batches pass straight from the warehouse adapter to stdout with no row materialization — pipe them into DuckDB, Polars, or pyarrow with zero JSON parsing. Requires an Arrow-capable build (the default build, or anyexec-*-arrow/exec-flightsql/exec-duckdbfeature; a--no-default-featuresbuild without one rejects--format arrowat runtime).
# Stream results into DuckDB for further analysis
dosi query --model model.yaml --metrics revenue --group-by orders.status \
--execute --connection prod-ch --format arrow \
| duckdb -c "SELECT * FROM read_arrow('/dev/stdin')"
# Or into Polars
dosi query ... --execute --format arrow \
| python -c "import polars as pl,sys; print(pl.read_ipc_stream(sys.stdin.buffer))"
Executing against a warehouse¶
--execute runs the compiled SQL and prints the rows. Without a connection it
uses local DuckDB — in-process and Arrow-native by default (bundled
exec-duckdb); --db <file> points at a DuckDB file, otherwise it is
in-memory. With --connection <name> it targets that profile's warehouse and
dialect.
$ dosi query --model model.yaml --metrics revenue --group-by orders.status \
--execute --connection prod-sr
Connection profiles use the
Datus agent.yml datasources:
vocabulary. Point --connections at a full agent.yml (reads
services.datasources) or a standalone datasources: file. Without the flag,
the file is discovered in order: DOSI_CONNECTIONS env →
./dosi-connections.yaml → ./osi-connections.yaml →
~/.config/dosi/connections.yaml → ~/.config/osi/connections.yaml →
./conf/agent.yml → ~/.datus/conf/agent.yml — an existing Datus install
works with zero config, and the pre-rename osi- paths still resolve.
Secrets interpolate from the environment as ${VAR}.
datasources:
prod-sr:
type: starrocks
host: sr.internal
port: 9030
arrow_flight_port: 9408 # opt into Arrow Flight SQL (SR ≥3.5.1)
username: osi
password: ${SR_PASSWORD}
database: analytics
default: true # used by --execute without --connection
prod-ch:
type: clickhouse
uri: http://ch.internal:8123
username: default
database: analytics
Resolution rules:
--executewithout--connectionuses the profile markeddefault: true; if none is marked, it falls back to local DuckDB (--dbor in-memory). Adefault: trueentry that failed to parse (e.g. an unset${VAR}) warns on stderr rather than silently using empty DuckDB.--dialectwith--connectionmust agree with the profile's dialect, or the command errors — drop--dialectand let the profile decide.
Warehouse drivers are feature-gated if you build from source, so a
hand-rolled binary only carries what you ask for (published builds carry them
all). Build with the features you need (or exec-all):
| Feature | Engines | Result path |
|---|---|---|
exec-duckdb (default) |
DuckDB (in-process) | Arrow-native |
exec-mysql |
MySQL, TiDB, StarRocks, Doris (MySQL wire) | rows |
exec-postgres |
Postgres | rows |
exec-hologres |
Hologres (Postgres wire; implies exec-postgres) |
rows |
exec-gaussdb |
GaussDB / openGauss (native SHA256 auth driver) | rows |
exec-oracle |
Oracle Database (ODPI-C; Instant Client at runtime) | rows |
exec-http |
ClickHouse, Trino | rows |
exec-http-arrow |
ClickHouse FORMAT ArrowStream |
Arrow-native |
exec-flightsql |
StarRocks / Doris Arrow Flight SQL (arrow_flight_port:) |
Arrow-native |
exec-snowflake |
Snowflake (SQL API v2 + key-pair JWT) | rows |
exec-bigquery |
BigQuery (jobs.query REST + service-account JWT) |
rows |
Per-engine setup and the Arrow result-path status live in connectors.md and arrow.md.
See also¶
- semantics.md — the normative behavior contract (metric inference, joins, fan-out protection, time handling) the CLI compiles to.
- rest-api.md — the same capabilities over HTTP.
- extensions-guide.md — the optional Datus model
extensions (
join_type,fill_nulls_with,time_dimension,time_granularity,time,dataset), and datus-extensions.md for their normative contract and versioning. - connectors.md — warehouse connector setup.
- arrow.md — enabling Arrow result transfer and what gets faster.