Direct completion
- Basis
- Accepted delivery
- Strength
- Simple and auditable
- Caution
- Does not establish downstream impact
Draft 0.1 — 2 August 2026
Measurement supports decisions and reward allocation. It does not guarantee scientific certainty, future performance, or an independently audited financial result.
ANUKA should connect contribution to outcome more rigorously than ordinary feedback tools and more honestly than marketing case studies.
The system must answer:
The protocol distinguishes:
A participant performed an action.
Example: Contributor submitted onboarding copy.
The product accepted the contribution.
The contribution reached a declared environment or audience.
A metric changed after or alongside the release.
The contribution caused some or all of the observed change under a declared method.
The result created revenue, savings, avoided loss, or another monetized effect.
A release plus a rising metric does not automatically justify a causal or financial claim.
Every outcome-dependent opportunity MUST define a measurement plan before release.
Required fields:
A reasoned account without source-connected quantitative evidence.
The metric changed across a time boundary, but alternative explanations are not controlled.
The result is compared against a relevant cohort, historical pattern, matched segment, or synthetic baseline.
A/B or other controlled assignment supports a stronger causal inference.
The effect was reproduced, externally reviewed, or observed across multiple relevant contexts.
The evidence level must be displayed with the result.
Evidence indicator
Higher levels strengthen the inference; they do not erase uncertainty, confounders or the need for clear public disclosure.
A reasoned account without source-connected quantitative evidence.
A change across time, with alternative explanations still open.
A relevant cohort, segment, history or synthetic baseline informs the result.
Controlled assignment supports a stronger causal inference.
The effect is reproduced, independently reviewed or observed in multiple relevant contexts.
Preferred where technically, ethically, and commercially appropriate.
Useful for time-dependent systems where treatment alternates across periods.
Compares cohorts introduced at different times.
Maintains an untreated group for evaluation.
Evaluates change against pre-intervention trend.
Compares similar users or accounts when randomization is unavailable.
Uses structured interviews, usability tests, or review panels for outcomes not captured adequately by one metric.
Confirms completion, performance, reliability, security, or compliance criteria without claiming commercial causality.
Each experiment has one declared primary metric unless a justified multi-objective design exists.
Guardrails detect harm such as:
A primary win with a material guardrail failure is not a successful loop.
Every metric SHOULD have a machine-readable contract:
Changing a metric definition during an experiment requires a new version and disclosure.
Possible baselines include:
The baseline source and limitations must be visible.
Reward is based on accepted delivery, not downstream business outcome.
A reward is earned when a defined threshold is met.
A reward is calculated from estimated incremental value attributable to the intervention.
A declared rule allocates credit among idea, implementation, review, rollout, and measurement roles.
A reviewer panel assigns bounded credit using evidence and disclosed criteria.
A fixed delivery reward is combined with a capped outcome bonus.
The MVP SHOULD prefer fixed or capped models over uncapped revenue participation.
Attribution comparison
Choose a model whose evidence burden matches the claim and reward risk. The MVP defaults toward fixed or capped approaches.
A result may depend on:
Credit allocation must not imply that one participant created the entire company-level effect.
Recommended fields:
Possible value categories:
Additional revenue reasonably attributable to the intervention.
Revenue adjusted for direct variable cost.
Reduction in spend or labor under a documented baseline.
Estimated prevention of churn, fraud, downtime, penalties, or rework.
Improvement in conversion, payback, utilization, or runway.
Learning that enables or rejects a future investment.
Financial claims should state whether values are:
Reward model reference
Set currency, timing, caps, rounding, reversals, taxes, and dispute rules alongside every model.
Every formula must specify currency, timing, caps, rounding, reversals, taxes, and dispute rules.
ANUKA must preserve failed experiments.
Possible outcomes:
A contributor may still earn delivery and learning rewards when the hypothesis fails honestly.
Punishing all negative findings creates pressure to manipulate results.
The plan must address:
Material events are added to the evidence package.
Before attribution, the system should check:
AI may:
AI MUST NOT silently:
Material AI-generated analysis should include model, tools, inputs, and human review state.
A public improvement story SHOULD include:
Avoid headlines such as ANUKA increased revenue by 42% when the evidence supports only a before/after association.
Traction reports may include verified source data and experiment results, but must distinguish:
ANUKA verification is not an audit unless a qualified independent auditor performed the audit.
Reviewers and analysts disclose:
High-value outcome bonuses SHOULD receive independent or multi-party review.
An evidence package SHOULD preserve:
Sensitive raw data remains access-controlled.