Skip to content

Attribution analysis

Every dashboard raises the same question: why did this metric change? Answering it by hand — or letting an AI agent answer it — normally means a loop of exploratory metric queries: totals for two periods, then one grouped query after another, then joining and comparing the numbers without slipping on segments that appear or disappear between the periods.

Attribution analysis collapses that loop into one call. Give Dosi a metric, the candidate dimensions, and two date windows; the engine runs everything itself and returns a ranked, ready-to-quote decomposition:

  • which dimension best explains the change,
  • which segment drove it, and by how much (as a share of the total change),
  • for ratio metrics, whether the change came from structure or from rate — e.g. "average order value rose because completed orders' own AOV went 75 → 100, even though their share of orders fell".

The numbers are computed by the engine with the method that is exact for each metric's type, and they add up — every decomposition reconciles against the totals, so the explanation is checkable rather than improvised.

Available on every surface: the CLI (dosi attribute), REST (POST /v1/query/attribute), MCP (the attribute_metric tool), and Python (Engine.attribute).

Quick start

dosi --model model.yaml attribute \
  --metric revenue \
  --dimensions status,customers.region,products.category \
  --baseline 2024-01-01..2024-02-01 \
  --current  2024-02-01..2024-03-01 \
  --db warehouse.duckdb

Which dimensions?

--dimensions is required in practice but not syntactically: omit it and the engine refuses with dimensions_required and hands you the shortlist for that metric, best first, plus a ready-to-paste fragment.

$ dosi --model model.yaml attribute --metric revenue \
    --baseline 2024-01-01..2024-02-01 --current 2024-02-01..2024-03-01
error: dimensions_required: name the dimensions to analyze; `candidates` lists what
  `revenue` can be grouped by, best first
  candidates: orders.status, customers.region, products.category
  retry with: "dimensions": ["orders.status", "customers.region", "products.category"]

The engine deliberately does not pick for you. Which columns explain a change is a question about the business, and the ranking the analysis itself uses cannot answer it: concentration (max |Δsegment| / |Δtotal|) rises with cardinality, so an engine left to guess puts customer ids above regions — and pays one warehouse query per guess to do it. The same shortlist is available ahead of time from dosi list dimensions --metric <name> (is_dimension: true rows, time columns aside), and what feeds it — declared or inferred — is described under D-DIM.

Windows are half-open ISO ranges (START..END). Use --connection for a named warehouse profile, --where to scope the whole analysis, and --time-dimension when the default time routing is not what you want. A parameterized metric (D-PARAM) takes --param name=value — one value per parameter, never a list — and the result echoes the resolved bindings in comparison_metadata.params.

Over REST or Python the request is the same shape:

from dosi_engine import Engine

engine = Engine("model.yaml")
result = engine.attribute({
    "metric": "revenue",
    "dimensions": ["status", "customers.region"],
    "baseline": {"start": "2024-01-01", "end": "2024-02-01"},
    "current":  {"start": "2024-02-01", "end": "2024-03-01"},
}, db_path="warehouse.duckdb")

What comes back

{
  "metric": "avg_order_value",
  "strategy": "mix_shift",                  // how the engine decomposed it
  "total_change": {"baseline_value": 75.0, "current_value": 76.67,
                   "delta": 1.67, "pct_change": 2.22},
  "dimension_ranking": [                    // ranked root-cause candidates
    {"dimension": "status", "score": 6.0}
  ],
  "top_dimension_values": [                 // the biggest movers
    {"dimension": "status", "value": "cancelled",
     "delta": 10.0, "contribution_pct": 600.0,
     "segment_kind": "entered",             // this segment is new this period
     "drill_down": {"where_sql": "status = 'cancelled'"}},
    {"dimension": "status", "value": "completed",
     "delta": -8.33, "contribution_pct": -500.0,
     "mix_effect": -28.69, "rate_effect": 20.35,
     "baseline_rate": 75.0, "current_rate": 100.0,
     "drill_down": {"where_sql": "status = 'completed'"}}
  ],
  "factor_totals": {"mix_effect": -18.7, "rate_effect": 20.4, ...},
  "warnings": []                            // structured caveats, see below
}

Reading it:

  • dimension_ranking orders the candidate dimensions by how well each one explains the change — read the first entry as "slice it this way".
  • contribution_pct is directly quotable: "this segment explains X% of the change". Percentages above 100 (or negative) mean segments moved in opposite directions and partly cancelled out — the response flags this.
  • segment_kind tells you when a segment is entered (new this period) or exited (gone this period) rather than a normal shift — the cases that silently skew hand-rolled comparisons. A fourth value, fallback, means the mix/rate split is undefined for that segment (zero or negative components); it still carries its whole delta, so the decomposition keeps adding up.
  • For ratio metrics, each segment's contribution splits into a mix effect (its weight in the population shifted) and a rate effect (its own ratio changed), with the per-segment rates and shares included — the "structure vs performance" narrative reads straight off the response, and the parts always sum to the total change. mix_effect, rate_effect, baseline_share and current_share are present only on normal segments: outside that path the underlying quantities are undefined, and the engine omits the field rather than publishing a number that does not mean what its name says. Check segment_kind (or simply the presence of mix_effect) before reading a share.
  • per_dimension (not shown) carries the full value-level detail per dimension, including whether that dimension's numbers reconcile with the totals.

How it helps an agent reason

The response is designed so an agent can go from "metric moved" to a verified root cause in one or two calls, without doing arithmetic:

  1. No probing loop. The ranking and contributions arrive pre-computed and pre-sorted; the agent quotes them instead of orchestrating and reconciling its own query sequence.
  2. Every finding carries its next step. Each segment includes drill_down.where_sql — a ready-to-paste filter. To go deeper, the agent calls attribute again with that filter (root-cause recursion), or hands it to a metric query to chart that segment's trend.
  3. Caveats are machine-readable. warnings is a list of stable codes — offsetting segments, truncated high-cardinality dimensions, a dimension whose numbers don't reconcile, near-zero total change — so the agent knows which conclusions to soften without parsing prose.
  4. Refusals guide instead of failing. A metric the engine cannot decompose faithfully (window metrics, distinct counts, complex expressions) returns strategy: "unsupported" with a structured reason, so the agent pivots — for a window metric, attribute its base metric over explicit windows — rather than retry-looping. dosi list metrics reports each metric's attribution_strategy up front, so support can be checked before asking.

Metric coverage

  • Additive metrics (term_wise) — sums, counts, and their linear combinations (including linear derived metrics, which additionally get a per-member breakdown of the change): full dimension attribution. An additive constant (SUM(x) + 100) no longer refuses: the variable part decomposes and the constant is reported in affine_constant.
  • Ratio metrics (mix_shift) — a single ratio of such linear forms, including averages: the structure/rate decomposition shown above.
  • General expressions (factor_shapley) — products of aggregates, nested ratios, and any other + − × ÷ arithmetic over SUM/COUNT measures: the engine runs one flat Shapley decomposition over the expression's leaf factors and allocates each factor's exact effect to segments in proportion to their share of that factor's change. The response adds a factors block (per-factor baseline/current/effect, summing to the total change exactly), and each segment's delta is its allocated share of the change — for these metrics baseline_value/current_value are the segment's own metric levels (context), not additive terms. Interaction terms split evenly between factors (the Shapley convention; for a product X·Y this is the familiar ΔX·Ȳ + ΔY·X̄ midpoint form). One deliberate exception: a plain ratio keeps the mix_shift structure/rate semantics rather than switching to Shapley — the two conventions differ numerically, and mix_effect/rate_effect are an established narrative.
  • Filter-derived metrics — attribute normally, and additionally carry filter_breakdown: the exact subset identity Δbase = Δfiltered + Δcomplement with the complement synthesized by the engine.
  • Offset window metrics (delta / percent_change / period-over- period) — the analysis decomposes the metric's base aggregate between your explicit windows and labels the transform in window_mapping; the transform itself is a per-bucket series, not a two-window comparison.
  • Not supported — frame/rank/value window metrics and distinct counts return a structured unsupported response instead of a misleading number. dosi list metrics reports each metric's attribution_strategy and an attributable flag (algebra + a usable default time axis) up front.
  • Levels of detail (D-LOD) — not yet. A metric that reads one (an aggregate over a fixed field, an include or exclude, or a compose over one) returns unsupported with reason level_of_detail naming what it reads, and is attributable: false up front; a candidate dimension that reads one is skipped with a dimension_skipped warning before any statement runs, and a where over one is refused. The reason is structural: a rollup is a function of each window's filter context, and the two-window statement computes it once — decomposing across a level of detail needs a rollup per window (see the D-LOD boundaries). To explain a share (revenue / revenue_grand), attribute the base metric: its mix_effect is the share's movement.

Limits

  • Up to 16 candidate dimensions per call, each analyzed independently; go deeper by recursing with drill_down.where_sql.
  • Per-dimension values are capped (default 500, max 1000); beyond the cap the dimension is flagged truncated. Under the cap the engine keeps one consistent top-N across both windows, ranked by combined magnitude — a segment that collapsed (or exploded) between the windows survives the cap as a whole pair, never as a half-seen phantom.
  • Segments present in only one window are handled for you — zero-filled or reported as entered/exited — so the comparison never silently drops them.
  • The engine issues one statement per dimension (plus one for the totals), each covering both windows at once — the analysis is consistent even on live-ingest warehouses, and a D-dimension call costs 1 + D round trips.

Errors

Two failure channels, deliberately distinct:

  • Bad request — the request itself is malformed (unknown metric, more than 16 dimensions, an invalid window): HTTP 400 / a CLI usage error. dimensions_required — no dimensions named — is one of these, and carries the candidate list; nothing reaches the warehouse.
  • Not computable — the request is fine but the data cannot support the analysis: denominator_nonpositive (a ratio metric whose denominator is zero or negative over a window) or nonfinite_evaluation (a factor_shapley expression hits a zero divisor on the windows' factor totals). REST returns HTTP 422 with {"error": {"code": "denominator_nonpositive", ...}}, Python raises NotComputableError, and the CLI prints the same structured code — retrying with different windows or filters may succeed, which is exactly what distinguishes it from a bad request.