Glossary

Overall evaluation criterion (OEC)

The single metric or composite formula a team agrees in advance will determine whether an experiment's treatment is judged a success.

An overall evaluation criterion is the metric, or a small weighted combination of metrics, that a team commits to before an experiment runs as the definition of success, so the result cannot be re-interpreted afterward by whichever metric happened to move favorably. It plays the same role in experimentation that a north star metric plays for a product overall, but scoped to a single test.

A simple OEC is one metric, such as seven-day retention or revenue per user. A composite OEC combines several signals into one score, for instance weighting short-term revenue against a measure of long-term engagement, when neither alone captures what a team actually cares about. Either way, the OEC is fixed in the experiment's pre-registration alongside its minimum detectable effect and sample size, precisely so a launch decision cannot be justified by scanning dozens of metrics after the fact and picking whichever reached significance.

The OEC matters because it forces hard tradeoffs, between short-term and long-term value, or between two competing metrics, to be made explicitly before the data arrives rather than implicitly afterward. A test with no agreed OEC tends to produce disputes about what "won," and is vulnerable to the same false-positive inflation that motivates false-discovery-rate correction when many metrics are checked.

Last reviewed September 22, 2026

In the index now

Related terms

Related guides