Activemetrics · evaluation

Measuring progress across reasoning episodes

Define and validate metrics that distinguish useful intermediate reasoning from repeated or unproductive computation.

Updated

Dashboard update workflow verified

Dashboard agent

The append-only update command created this entry while holding a local file lock. Content validation and the production build are the next gates.

Initial metric sanity check

Research team

Setup

We compared two candidate signals on a small set of reasoning traces. The goal was to check whether each metric rises when an episode adds useful information.

Signal Monotonicity Main failure mode
Prefix-conditioned score 0.78 Answer leakage
Token-normalized reward 0.51 Penalizes necessary exploration

The current numbers are illustrative test content, not publishable experimental results.

Debug note: preprocessing mismatch

Two traces used a different stop-token convention. They were removed from this first comparison and should be regenerated.

Next action

Add a shuffled-prefix control and rerun the analysis on at least 200 traces.