Attribution analysis¶
Every dashboard raises the same question: why did this metric change? Answering it by hand — or letting an AI agent answer it — normally means a loop of exploratory metric queries: totals for two periods, then one grouped query after another, then joining and comparing the numbers without slipping on segments that appear or disappear between the periods.
Attribution analysis collapses that loop into one call. Give Dosi a metric, the candidate dimensions, and two date windows; the engine runs everything itself and returns a ranked, ready-to-quote decomposition:
- which dimension best explains the change,
- which segment drove it, and by how much (as a share of the total change),
- for ratio metrics, whether the change came from structure or from rate — e.g. "average order value rose because completed orders' own AOV went 75 → 100, even though their share of orders fell".
The numbers are computed by the engine with the method that is exact for each metric's type, and they add up — every decomposition reconciles against the totals, so the explanation is checkable rather than improvised.
Available on every surface: the CLI (dosi attribute), REST
(POST /v1/query/attribute), MCP (the attribute_metric tool), and Python
(Engine.attribute).
Quick start¶
dosi --model model.yaml attribute \
--metric revenue \
--dimensions status,customers.region,products.category \
--baseline 2024-01-01..2024-02-01 \
--current 2024-02-01..2024-03-01 \
--db warehouse.duckdb
Which dimensions?¶
--dimensions is required in practice but not syntactically: omit it and the
engine refuses with dimensions_required and hands you the shortlist for that
metric, best first, plus a ready-to-paste fragment.
$ dosi --model model.yaml attribute --metric revenue \
--baseline 2024-01-01..2024-02-01 --current 2024-02-01..2024-03-01
error: dimensions_required: name the dimensions to analyze; `candidates` lists what
`revenue` can be grouped by, best first
candidates: orders.status, customers.region, products.category
retry with: "dimensions": ["orders.status", "customers.region", "products.category"]
The engine deliberately does not pick for you. Which columns explain a change
is a question about the business, and the ranking the analysis itself uses
cannot answer it: concentration (max |Δsegment| / |Δtotal|) rises with
cardinality, so an engine left to guess puts customer ids above regions — and
pays one warehouse query per guess to do it. The same shortlist is available
ahead of time from dosi list dimensions --metric <name> (is_dimension:
true rows, time columns aside), and what feeds it — declared or inferred — is
described under D-DIM.
Windows are half-open ISO ranges (START..END). Use --connection for a
named warehouse profile, --where to scope the whole analysis, and
--time-dimension when the default time routing is not what you want. A
parameterized metric (D-PARAM) takes
--param name=value — one value per parameter, never a list — and the
result echoes the resolved bindings in comparison_metadata.params.
Over REST or Python the request is the same shape:
from dosi_engine import Engine
engine = Engine("model.yaml")
result = engine.attribute({
"metric": "revenue",
"dimensions": ["status", "customers.region"],
"baseline": {"start": "2024-01-01", "end": "2024-02-01"},
"current": {"start": "2024-02-01", "end": "2024-03-01"},
}, db_path="warehouse.duckdb")
What comes back¶
{
"metric": "avg_order_value",
"strategy": "mix_shift", // how the engine decomposed it
"total_change": {"baseline_value": 75.0, "current_value": 76.67,
"delta": 1.67, "pct_change": 2.22},
"dimension_ranking": [ // ranked root-cause candidates
{"dimension": "status", "score": 6.0}
],
"top_dimension_values": [ // the biggest movers
{"dimension": "status", "value": "cancelled",
"delta": 10.0, "contribution_pct": 600.0,
"segment_kind": "entered", // this segment is new this period
"drill_down": {"where_sql": "status = 'cancelled'"}},
{"dimension": "status", "value": "completed",
"delta": -8.33, "contribution_pct": -500.0,
"mix_effect": -28.69, "rate_effect": 20.35,
"baseline_rate": 75.0, "current_rate": 100.0,
"drill_down": {"where_sql": "status = 'completed'"}}
],
"factor_totals": {"mix_effect": -18.7, "rate_effect": 20.4, ...},
"warnings": [] // structured caveats, see below
}
Reading it:
dimension_rankingorders the candidate dimensions by how well each one explains the change — read the first entry as "slice it this way".contribution_pctis directly quotable: "this segment explains X% of the change". Percentages above 100 (or negative) mean segments moved in opposite directions and partly cancelled out — the response flags this.segment_kindtells you when a segment isentered(new this period) orexited(gone this period) rather than a normal shift — the cases that silently skew hand-rolled comparisons. A fourth value,fallback, means the mix/rate split is undefined for that segment (zero or negative components); it still carries its wholedelta, so the decomposition keeps adding up.- For ratio metrics, each segment's contribution splits into a
mix effect (its weight in the population shifted) and a rate
effect (its own ratio changed), with the per-segment rates and shares
included — the "structure vs performance" narrative reads straight off
the response, and the parts always sum to the total change.
mix_effect,rate_effect,baseline_shareandcurrent_shareare present only onnormalsegments: outside that path the underlying quantities are undefined, and the engine omits the field rather than publishing a number that does not mean what its name says. Checksegment_kind(or simply the presence ofmix_effect) before reading a share. per_dimension(not shown) carries the full value-level detail per dimension, including whether that dimension's numbers reconcile with the totals.
How it helps an agent reason¶
The response is designed so an agent can go from "metric moved" to a verified root cause in one or two calls, without doing arithmetic:
- No probing loop. The ranking and contributions arrive pre-computed and pre-sorted; the agent quotes them instead of orchestrating and reconciling its own query sequence.
- Every finding carries its next step. Each segment includes
drill_down.where_sql— a ready-to-paste filter. To go deeper, the agent callsattributeagain with that filter (root-cause recursion), or hands it to a metric query to chart that segment's trend. - Caveats are machine-readable.
warningsis a list of stable codes — offsetting segments, truncated high-cardinality dimensions, a dimension whose numbers don't reconcile, near-zero total change — so the agent knows which conclusions to soften without parsing prose. - Refusals guide instead of failing. A metric the engine cannot
decompose faithfully (window metrics, distinct counts, complex
expressions) returns
strategy: "unsupported"with a structured reason, so the agent pivots — for a window metric, attribute its base metric over explicit windows — rather than retry-looping.dosi list metricsreports each metric'sattribution_strategyup front, so support can be checked before asking.
Metric coverage¶
- Additive metrics (
term_wise) — sums, counts, and their linear combinations (including linear derived metrics, which additionally get a per-member breakdown of the change): full dimension attribution. An additive constant (SUM(x) + 100) no longer refuses: the variable part decomposes and the constant is reported inaffine_constant. - Ratio metrics (
mix_shift) — a single ratio of such linear forms, including averages: the structure/rate decomposition shown above. - General expressions (
factor_shapley) — products of aggregates, nested ratios, and any other+ − × ÷arithmetic over SUM/COUNT measures: the engine runs one flat Shapley decomposition over the expression's leaf factors and allocates each factor's exact effect to segments in proportion to their share of that factor's change. The response adds afactorsblock (per-factor baseline/current/effect, summing to the total change exactly), and each segment'sdeltais its allocated share of the change — for these metricsbaseline_value/current_valueare the segment's own metric levels (context), not additive terms. Interaction terms split evenly between factors (the Shapley convention; for a productX·Ythis is the familiarΔX·Ȳ + ΔY·X̄midpoint form). One deliberate exception: a plain ratio keeps themix_shiftstructure/rate semantics rather than switching to Shapley — the two conventions differ numerically, andmix_effect/rate_effectare an established narrative. - Filter-derived metrics — attribute normally, and additionally carry
filter_breakdown: the exact subset identityΔbase = Δfiltered + Δcomplementwith the complement synthesized by the engine. - Offset window metrics (
delta/percent_change/ period-over- period) — the analysis decomposes the metric's base aggregate between your explicit windows and labels the transform inwindow_mapping; the transform itself is a per-bucket series, not a two-window comparison. - Not supported — frame/rank/value window metrics and distinct counts
return a structured
unsupportedresponse instead of a misleading number.dosi list metricsreports each metric'sattribution_strategyand anattributableflag (algebra + a usable default time axis) up front. - Levels of detail (D-LOD) — not yet. A
metric that reads one (an aggregate over a
fixedfield, anincludeorexclude, or a compose over one) returnsunsupportedwith reasonlevel_of_detailnaming what it reads, and isattributable: falseup front; a candidate dimension that reads one is skipped with adimension_skippedwarning before any statement runs, and awhereover one is refused. The reason is structural: a rollup is a function of each window's filter context, and the two-window statement computes it once — decomposing across a level of detail needs a rollup per window (see the D-LOD boundaries). To explain a share (revenue / revenue_grand), attribute the base metric: itsmix_effectis the share's movement.
Limits¶
- Up to 16 candidate dimensions per call, each analyzed independently; go
deeper by recursing with
drill_down.where_sql. - Per-dimension values are capped (default 500, max 1000); beyond the cap
the dimension is flagged
truncated. Under the cap the engine keeps one consistent top-N across both windows, ranked by combined magnitude — a segment that collapsed (or exploded) between the windows survives the cap as a whole pair, never as a half-seen phantom. - Segments present in only one window are handled for you — zero-filled or reported as entered/exited — so the comparison never silently drops them.
- The engine issues one statement per dimension (plus one for the
totals), each covering both windows at once — the analysis is consistent
even on live-ingest warehouses, and a D-dimension call costs
1 + Dround trips.
Errors¶
Two failure channels, deliberately distinct:
- Bad request — the request itself is malformed (unknown metric, more
than 16 dimensions, an invalid window): HTTP 400 / a CLI usage error.
dimensions_required— no dimensions named — is one of these, and carries the candidate list; nothing reaches the warehouse. - Not computable — the request is fine but the data cannot support
the analysis:
denominator_nonpositive(a ratio metric whose denominator is zero or negative over a window) ornonfinite_evaluation(afactor_shapleyexpression hits a zero divisor on the windows' factor totals). REST returns HTTP 422 with{"error": {"code": "denominator_nonpositive", ...}}, Python raisesNotComputableError, and the CLI prints the same structured code — retrying with different windows or filters may succeed, which is exactly what distinguishes it from a bad request.