The dosi-engine semantics contract¶
The Apache Ossie (formerly OSI) core spec deliberately leaves execution semantics implicit: a metric is a name plus a raw SQL string, a relationship is a column mapping, and a field can declare a logical type and a time-dimension role. This document is the normative contract for how dosi-engine fills those gaps. It is versioned with the engine; behavioral changes here are breaking changes.
Rules are labeled S-<area>-<n> for reference from issues and tests.
Fields and metrics accept the optional Ossie core datatype: String, Integer,
Decimal, Float, Boolean, Date, Time, DateTime, DateTimeTz, or Opaque
(case-sensitive). It describes the expression's result, not necessarily the
source column. Omit it when unknown; Opaque means a known type outside the
portable vocabulary. Decimal leaves precision and scale unspecified;
DateTime has no timezone, while DateTimeTz identifies an instant with
offset/timezone context. The declaration does not insert casts, change metric
classification, or set output display formats. Both engine modes accept it.
1. Expressions¶
- S-EXPR-1 Every
Expressionis compiled from itsANSI_SQLdialect entry when present, else the first SQL-family entry (SNOWFLAKE,DATABRICKS).MDX/TABLEAU/MAQLentries are ignored; an expression with only non-SQL entries is a compile error for metrics and a warning for fields. - S-EXPR-2 Expression fragments are parsed as a single scalar expression. Anything else — multiple projections, trailing aliases, statements — is a parse error (this is also the injection barrier: user text never reaches SQL by string concatenation anywhere in the engine).
2. Metric inference¶
- S-METRIC-1 A metric expression must contain at least one aggregate
call. Supported aggregates:
SUM,COUNT,COUNT(DISTINCT x),AVG,MIN,MAX. Any other aggregate function isunsupported_aggregate. - S-METRIC-2 Classification, applied to the parenthesis-stripped root:
- a single aggregate call → aggregate metric;
agg / agg(both sides bare aggregates) → ratio; the division is wrappedCAST(numerator AS DOUBLE) / denominatorto avoid integer division on truncating engines;- anything else → expression; each aggregate subtree becomes a
measure, the surrounding arithmetic is preserved verbatim (including
NULLIF,CASE, …). NoteSUM(a)/NULLIF(SUM(b),0)is an expression, not a ratio, and is not double-cast — the authored SQL wins. - S-METRIC-3 Rejected in metric expressions, as structured errors:
window functions (
window_in_metric), subqueries (subquery_in_metric), nested aggregates (nested_aggregate), column references outside any aggregate (bare_column_in_metric), aggregates over no column (SUM(1)→bare_column_in_metric), multi-columnCOUNT(DISTINCT a, b), and references to other metrics — metric references live exclusively in the DATUSderiveextension key (docs/datus-extensions.md#d-derive), never in expression SQL; theexpressionof a derived metric stays a self-contained flattened form.
3. Measures (synthesized)¶
- S-MEASURE-1 Each distinct aggregate call becomes a measure named
{dataset}_{stem}_{suffix}: stem = the aggregate argument rendered and sanitized (lowercase, non-alphanumeric runs →_); suffix =sum,count,count_distinct,average,min,max;COUNT(*)→{dataset}_rows_count. - S-MEASURE-2 Measures dedupe by signature (dataset + aggregate kind +
distinct + normalized argument), within and across metrics. Two different
aggregates that sanitize to the same name are a compile error
(
measure_name_collision) — never silently merged.
4. Column → dataset attribution¶
- S-ATTR-1
dataset.columnresolves exactly; the qualifier must be a dataset name and the column one of its declared fields. Metric and filter columns reference fields (which may themselves be expressions), not raw physical columns. - S-ATTR-2 A bare column resolves iff exactly one dataset declares a
field with that name; zero →
unknown_column, several →ambiguous_column(with candidates). - S-ATTR-3 All columns inside one aggregate must belong to one dataset
(
cross_dataset_aggregateotherwise). - S-ATTR-4
COUNT(*)attributes to the unique dataset of the metric's other measures, or to the model's only dataset; otherwisecount_star_needs_dataset.
5. Joins¶
- S-JOIN-1 Relationships are the only join source. Edges are followed
strictly many→one (
from→to), so walking a path never multiplies the origin's rows. Joins are on the AND-ed equality of the relationship's column pairs, and areLEFT JOINby default (orphan many-side rows survive into a NULL dimension bucket). A relationship may declareINNERvia the Datusjoin_typeextension (drop orphans — attribution semantics; docs/datus-extensions.md#d-join). - S-JOIN-2 A target dataset is reachable iff exactly one simple path
exists (edge-level: parallel relationships are distinct paths). Zero paths
→
no_join_path; several →ambiguous_join_pathwith each candidate path spelled out. Maximum path length: 6 hops. - S-JOIN-3 Self-relationships and cyclic walks are never followed.
- S-JOIN-4 A relationship declared
cardinality: many_to_many(Datus D-CONFORM, docs/datus-extensions.md#d-conform) is not a join edge: the graph never follows it in either direction, and naming it in a relationship path isinvalid_relationship_path. Its column pairs declare one dimension spelled on two datasets — the transitive closure of every such pair is a conformed class — and are consumed by S-FAN-7 instead.
6. Fan-out protection & branch assignment¶
The correctness core. Terms: a measure's home is the dataset its columns live on; the required set of a query is every dataset referenced by group-by dims and filters.
- S-FAN-1 Candidate evaluation bases for a measure: datasets that reach
(S-JOIN-2) both the measure's home and the entire required set.
Duplicate-sensitive aggregates (
SUM,AVG,COUNT) are additionally restricted to their home — evaluating them over a fanned-out join would double-count.COUNT DISTINCT,MIN,MAXmay also evaluate from any dataset that reaches their home. - S-FAN-2 A measure evaluates at its home whenever the home is a
candidate. The dataset an aggregate names is its row set, so
COUNT(DISTINCT dim.k)counts everydimentity in the group whether it stands alone or shares a ratio with a fact measure — one definition, one number, regardless of what else is in the query. Only a measure whose home cannot reach the required set moves: groupingSUM(fact.x) / COUNT(DISTINCT dim.k)by a dimension that onlyfactreaches evaluates the count over the fact join — on the entities the fact observes — because nothing else can serve that grouping. A measure evaluated away from home is projected as<measure>__at_<base>, so two placements of one measure in a query never share a column. To count observed entities under every grouping, aggregate the fact's foreign key instead:COUNT(DISTINCT fact.dim_fk)lives on the fact and never moves. A time range ormetric_timeadds the fact's time column to the required set, so a time-scoped ratio's dimension count follows the fact too. - S-FAN-3 Measures on different bases aggregate in separate branches,
each at its base's grain. Branches all produce the requested dim columns and merge
FULL OUTER JOINon them; the k-th branch joins on the null-safe equality ofCOALESCE(m0.key … m(k-1).key)andmk.key(see S-FAN-5), and final keys project asCOALESCEacross branches. Zero group-by keys →CROSS JOIN. Where a dialect cannot express that join — the Postgres family requires an equi condition in a FULL JOIN, MySQL/TiDB have no FULL JOIN, and ClickHouse cannot use aCOALESCEas a join key — the same merge is emitted as akeysCTE (UNIONof every branch's key tuples) that each branchLEFT JOINs back. Both shapes yield identical rows; the choice is invisible in the result and visible in--explain/generated SQL only. - S-FAN-4 A duplicate-sensitive measure whose home has no candidate base
is a
fan_out_riskerror, with a retry hint. The engine never silently emits a double-counting query. - S-FAN-5 A group-key value that is NULL merges as one row across branches,
not one per branch: the merge join is null-safe, rendered as the portable
(a = b OR (a IS NULL AND b IS NULL))(a bare=would leave each branch's NULL bucket unmatched — SQLNULL = NULLis unknown). ClickHouse rejects that OR-expansion as a join key, so there — and only there — it is emitted as the equivalent single predicateIS NOT DISTINCT FROM. - S-FAN-6 A bare count metric (its value is exactly one COUNT / COUNT
DISTINCT measure) reads 0, not NULL, for a group present only in other
branches — a count over no rows is 0. This applies only when the whole metric
is the count; a count embedded in a ratio or expression keeps NULL so it
propagates (and a count denominator never becomes a literal
0divisor). A metric's Datusfill_nulls_withextension overrides this default and applies to any metric kind (docs/datus-extensions.md#d-fill). - S-FAN-7 Conformed resolution is per branch. A group-by dimension, WHERE
column or time-range column whose dataset a branch's base does not reach
along many→one edges resolves to the member of its conformed class (S-JOIN-4)
that lives on the base — or the one member the base reaches — and projects
under the query's output name, so the branches still merge on one column
(S-FAN-3).
metric_timeresolves the same way when the facts' primary time dimensions form one class: each branch filters and groups its own date. A grain truncates the resolved column, under that column's own D-GRAIN floor. Only paired columns resolve this way; an unpaired column keeps the ordinaryunconformed_dimension/no_join_path. A model with no conformed class plans exactly as before.
7. Dimensions, grains, time¶
- S-TIME-1 Group-by items are
dataset.fieldor a bare unique field name; output columns are named{field}, or{field}__{grain}when a query-time grain is applied. Duplicate output names are an error. - S-TIME-2 Grains (
day,week,month,quarter,year) apply only to fields whose effectiveis_timeis true, lowering toDATE_TRUNC(per-dialect idioms where the dialect lacks it: MySQL/TiDB useDATE()/STR_TO_DATE(DATE_FORMAT(...))/YEARWEEKrewrites, ISO weeks starting Monday). An explicitdimension.is_timetakes precedence; otherwiseDate,Time,DateTime, andDateTimeTzdefault to true, including whendimensionis absent. Other or omitted datatypes default to false. Useis_time: falseto exclude audit timestamps. This role rule applies in both modes. AcceptingTimemetadata does not add time-of-day ranges or sub-day grains: current time queries require date-bearing values. - S-TIME-3 A time range is half-open
[start, end)over ISO dates, compiled asfield >= DATE start AND field < DATE endbefore aggregation, against: the explicitly named time dimension, else the single time item in the group-by (metric_timecounts as one), else — with no time item at all — each metric's primary time dimension (S-TIME-5). Only a group-by holding several time items still needs an explicit name (time_range_needs_dimension). - S-TIME-4 Ossie core carries no granularity metadata; native grain is
whatever the field expression yields, unless the field declares one via
the Datus
time_granularityextension (docs/datus-extensions.md#d-grain) — then requesting a strictly finer grain is agrain_too_fineerror. Coredatatypeidentifies the logical type but not a string or integer date encoding ('20260101',202601). These require the Datustimeextension (docs/datus-extensions.md#d-format) to describe their layout;String/Integeralone cannot select a parser. The encoding also supplies the native grain whentime_granularityis absent. A time field withdatatype: Datehas native day grain without a Datus declaration. Window metrics (period-over-period / rolling / cumulative) are the Datus D-WINDOW extension consuming the S-TIME-5 axis — see window-extension.md; a time spine remains proposed upstream. - S-TIME-5 Every metric may have a primary (aggregation) time
dimension, resolved as: the metric-level Datus
time_dimension(docs/datus-extensions.md#d-time), else the unique primary time among the metric's datasets — a dataset's being its explicittime_dimensionextension, else its singleis_timefield (this last inference reads no extension and applies in basic mode too). The reserved query namemetric_timegroups/filters by it: in a multi-branch plan each branch substitutes its own primary time column under the shared output name (metric_time/metric_time__{grain}) and branches merge on that name, so unrelated facts align on their respective business time axes. A metric with no resolvable primary time isno_primary_time_dimension. Two metrics on one base dataset with different primary times — a compose on a conformed month beside a member on its own fact's month — each get their own branch over that base, filtered and grouped on their own axis, and merge on the group keys like any two branches: every metric reads exactly what it reads alone, at the cost of one more scan of that base per axis. A model field literally namedmetric_timeis shadowed by the reserved name — qualify it asdataset.metric_time.
8. Filters¶
- S-FILTER-1
--where(where_sql) is scalar boolean SQL. It is split into its top-levelANDconjuncts, and each conjunct runs in the position that can evaluate it — decided by which columns exist there, never by guessing intent: - a conjunct that names a selected metric filters the result rows:
after every metric, window level and compose is computed, before
order_by/limit— so "the top 10 of the groups above X" means what it says. Undertime_rangesthe name is{metric}__{tag}; - in a window query, a conjunct that names only group-by fields also filters the result, so the rows kept keep the ranks, partition counts and weights the whole population gave them;
- every other conjunct — one naming a field the query does not group by, or a time field grouped at a grain — filters the input rows before aggregation, in every branch. Columns resolve per S-ATTR rules; filter datasets join into each branch like dims. When a branch carries a fixed level-of-detail rollup (D-LOD), such a conjunct is a dimension filter: it applies after the rollup is computed and never scopes it.
Outside window queries a condition on group-by fields stays on the input:
filtering a grouping key before or after aggregating gives the same rows
and the same values, and the input position prunes the scan. A reference
resolves as a field first, so a metric sharing a field's name does not move
an existing condition.
- S-FILTER-2 context_filter is row-level SQL that always filters the
input — the population a window ranks, counts and weighs. The same
predicate answers a different question in each place:
where_sql: "market = 'North'" shows North's rank among all markets,
context_filter: "market = 'North'" ranks North against itself. The time
range (time_range / time_ranges) sits in this position too; a window's
lookback scan widens it and trims back to the requested bounds after the
window, so it is a calculation window, not a display selection. Both are
the context a fixed level-of-detail rollup is computed over, so neither may
read a level-of-detail value (unsupported_filter).
- S-FILTER-3 Refused: subqueries, window functions and bind parameters
(unsupported_filter); inline aggregates (aggregate_in_where — declare
the metric and name it instead); a metric the query does not select
(result_filter_metric_not_selected); one conjunct that names both a
metric and an ungrouped field (result_filter_unprojected — it has no
position, so group by the field or split the conjunct).
explain shows each position: ResultFilter[…] above the projection root,
Filter on the scan. A result filter lowers to one more SELECT over the whole
query (WITH …, result_rows AS (…) SELECT … FROM result_rows WHERE …),
because the outermost window level computes in its own projection, where a
WHERE would run first.
9. Datasets¶
- S-DATA-1 A
sourcecontaining whitespace is an inline query (compiled as a derived table); otherwise it is a table reference split on.into up to catalog.schema.table. Identifier quoting is not yet supported. - S-DATA-2 Within a branch, each dataset appears at most once and is aliased by its dataset name (no self-joins in v1).
10. Field references & explicit paths¶
Both query planes resolve a field reference the same way, in one place.
- S-VIA-1 A reference is
field,dataset.field, orrelationship[.relationship…].field. The first segment is matched against dataset names first, so a model whose relationship shares a dataset's name keeps resolving to the dataset — adding this spelling cannot change what an existing reference means. - S-VIA-2 A relationship-prefixed reference names its own join path
edge by edge. The edges must chain (each leaves the dataset the previous
one arrived at), must not revisit a dataset, and are capped at the graph's
6-hop bound. Otherwise
unknown_relationship/invalid_relationship_path. - S-VIA-3 In a metric query the branch base is chosen after resolution,
so a named path pins it: the base must be the path's first edge's
fromdataset. Two named paths that disagree on their origin are an error, as is a path whose origin cannot host the query's measures. - S-VIA-4 Filters accept all three spellings, at any path length. A reference of three or more parts is not a SQL column — it parses as struct-field access — so the planner resolves those separately from ordinary columns. Every reference in a filter is resolved against the model or the filter is refused; none reaches the generated SQL unvalidated.
- S-VIA-5 A query reaches each dataset by exactly one path. Lowering
aliases tables by dataset name (S-DATA-2), so two paths to one dataset
would read the same alias twice; that is
not_implemented, never silently-equal columns.
11. Detail queries (projection plane)¶
- S-SELECT-1
fromfixes the result grain: one output row per row of that dataset. Every other dataset is reached along many→one edges only (S-JOIN-1), so no join in the path can multiply the root's rows. A dataset reachable only against that direction isdetail_fanout, whosesuggested_retrynames the root that answers the question at the finer grain. - S-SELECT-2 The required set is every dataset named by a projection, a filter column, or the time range. All of them are joined; only projections become output columns.
- S-SELECT-3 An output column is the reference with the
fromprefix stripped and each remaining.replaced by__, plus__{grain}when a grain is applied. Root fields stay bare; everything else carries its path. Two fields yielding one name isduplicate_output_name. - S-SELECT-4 A detail query projects row-level fields only. A metric
name in the projection is
metric_in_detail_query; an aggregate inwhereisaggregate_in_where(S-FILTER-2). - S-SELECT-5 With no
dimension, a time range routes to the root dataset's primary time dimension (D-TIME), else to the single time field among the projections, elsetime_range_needs_dimension. - S-SELECT-6 Order keys name projected output columns
(
unknown_order_key).limitdefaults to 100 and is clamped to 10000 by every surface.
12. Errors¶
- S-ERR-1 Every compile/query error carries a stable snake_case
code, a human message, the names involved,candidateswhere a bad reference has alternatives, andsuggested_retrywhere a rewrite would succeed.--format jsonemits the full structure. Error text is not a stable API; codes are.