Counting Changed Decisions Before Deploy with Decision Replay
Replay a day of stored decisions on the candidate version to count how many real decisions a rule change would move.
The problem
A marketplace auto-approves refunds up to $50 from established accounts, and the team wants to raise the limit to $75. Replaying one day of stored refund decisions on the $75 version moves 407 of that day's 2,000 refunds from review to automatic approval.
Refunds that are not auto-approved go to a support agent, and the review queue held 1,176 that day. At $75 it would have held 769, and automatic approvals would have risen from 824 to 1,231. Support asked for the change to shorten that queue. Finance reads the same 407 as refunds that would go out with no one looking at them.
Replaying the same day on the live $50 version, a control run, changes none of the 2,000 decisions. A control run would report changes if another version had decided part of the day, as happens around a deploy. So would failed calls in the day (see Edge cases). The control run's 0 is what lets the 407 stand for the new limit alone.
The refund router makes three checks. The amount must be within the limit, the account at least 30 days old, and its refunds in the last 90 days fewer than three. Unit tests can prove each check with hand-picked refunds. No test can say how many of a day's real refunds pass all three at $75 but not at $50.
Counting the day's refunds by amount alone gives 473. Why should support plan around 407?
The naive approach
Suppose the router still lived in code: one Java method with the three checks. Sizing a new limit would then mean writing a second method beside it.
public class RefundRouter {
static final BigDecimal LIMIT_USD = new BigDecimal("50.00");
static final int MIN_ACCOUNT_AGE_DAYS = 30;
static final int MAX_RECENT_REFUNDS = 3;
private final AccountRepository accounts;
private final RefundRepository refunds;
public Route route(RefundRequest req, Instant now) {
Account account = accounts.find(req.accountId());
long ageDays = Duration.between(account.createdAt(), now).toDays();
if (ageDays < MIN_ACCOUNT_AGE_DAYS) {
return Route.MANUAL_REVIEW;
}
Instant since = now.minus(Duration.ofDays(90));
if (refunds.countSince(req.accountId(), since) >= MAX_RECENT_REFUNDS) {
return Route.MANUAL_REVIEW;
}
if (req.amountUsd().compareTo(LIMIT_USD) <= 0) {
return Route.AUTO_APPROVE;
}
return Route.MANUAL_REVIEW;
}
// Sizing the $75 proposal: one day's refunds between the two limits.
public long estimateMoved(LocalDate day, BigDecimal newLimitUsd) {
return refunds.findCreatedOn(day).stream()
.filter(r -> r.amountUsd().compareTo(LIMIT_USD) > 0)
.filter(r -> r.amountUsd().compareTo(newLimitUsd) <= 0)
.count();
}
}
Run over the 2,000 refunds of 2026-10-07, estimateMoved would return 473: every refund above $50 and up to $75.
The estimate copies one of the router's three checks. The real router runs in LexQ, covered in the next section, and LexQ stored each refund's inputs. Counted from those, 66 of the 473 still go to review at $75. Fifty-three come from accounts younger than 30 days, and 13 from older accounts with three or more refunds in 90 days. The estimate overstates the change by 16%.
Adding the other two checks means rebuilding each refund's inputs as they stood when it arrived. In code, route computes the account's age and its 90-day refund count at that moment and stores neither. The estimate can rebuild both from timestamps, but that is a second copy of route's logic, written against history. It has to match route on details such as whether the 90-day count includes the refund being routed.
Nothing in code records the routes route returns. The estimate assumes the current code decided every refund that day. If a hotfix changed route at noon, the morning's refunds went through code that no longer exists, and the estimate cannot tell. Each later change to route also needs a matching edit to estimateMoved, which no test enforces.
Logging each request's inputs and route would end both problems: no inputs to rebuild, no guessing which code decided. One query would then count the same 407. That query still copies the plan. A person writes its conditions from "raise the limit to $75", so a mistake made while editing the rules never reaches it. Suppose the edit that raised the limit also turned the age check from 30 into 3 by mistake. The query still returns 407. Rerunning the day's refunds through the mistaken rules with the method in the next section changes 535 decisions, including refunds from accounts 3 days old.
Defining the pattern
Decision Replay runs stored production calls through a candidate version and compares each result with the decision stored when the call ran. LexQ, the hosted service that runs the router's rules, stores every call. The refunds service calls LexQ once per refund and sends facts, named values such as refundAmountUsd. The rules answer with actions, such as setting the fact refundRoute to AUTO_APPROVE, and LexQ stores both with the call.
The stored decision is the baseline, and it is never recomputed. A window replay reads one policy group's calls over a window of whole days, set with --from and --to. A policy group holds one decision's rules and their versions. The replay runs only the candidate on each stored input and counts the calls whose actions differ from the baseline.
Replay fits a policy group that already serves production calls, like the refund router, when the change is a new version of that group. A router that has not gone live has no stored calls. Impact Simulation gives its estimate instead by running both versions over an uploaded file of sample requests, as in testing a rule change before deploy. A candidate that reads a new fact does not fit either, because no stored call sent it.
In LexQ the router is one policy group with two rules. The auto-approve rule holds the three checks and sets refundRoute to AUTO_APPROVE. A catch-all, whose empty condition matches every refund, sets it to MANUAL_REVIEW. Both rules carry the mutex group key refund-route, so when both match a refund only the higher-priority rule, the one listed first, is selected. The live version, v1, has the $50 limit. The candidate, v2, is a draft cloned from v1, with the limit raised to 75 and the rule renamed to match.
{
"name": "Auto-approve: up to $75, established account",
"condition": {
"type": "GROUP",
"operator": "AND",
"children": [
{
"type": "SINGLE",
"field": "refundAmountUsd",
"operator": "LESS_THAN_OR_EQUAL",
"value": 75,
"valueType": "NUMBER"
},
{
"type": "SINGLE",
"field": "accountAgeDays",
"operator": "GREATER_THAN_OR_EQUAL",
"value": 30,
"valueType": "NUMBER"
},
{
"type": "SINGLE",
"field": "refundsLast90Days",
"operator": "LESS_THAN",
"value": 3,
"valueType": "NUMBER"
}
]
},
"actions": [
{
"type": "SET_FACT",
"parameters": { "targetVar": "refundRoute", "value": "AUTO_APPROVE" }
}
],
"mutexGroup": "refund-route",
"mutexMode": "EXCLUSIVE",
"mutexStrategy": "HIGHEST_PRIORITY",
"mutexLimit": 1,
"isEnabled": true
}
The refunds service now computes accountAgeDays and refundsLast90Days when the refund arrives and sends them with refundAmountUsd. The values the rules saw are therefore stored with the call, next to the route they produced.
Decision Replay strategy
The check is two replays of the same day. The first, the control run, uses the live v1 as its candidate. For 2026-10-07 it changes 0 of 2,000 decisions: v1 decided every stored call that day, and its rules reproduce the same actions. The second run uses v2, and with the control at 0, every change it counts comes from the new limit.
# Control: v1 against the decisions it made
lexq replay start --version-id <v1-version-id> \
--from 2026-10-07 --to 2026-10-07 --max-records 5000
# Candidate: v2 against the same decisions
lexq replay start --version-id <v2-version-id> \
--from 2026-10-07 --to 2026-10-07 --max-records 5000
--max-records is set above the day's 2,000 calls. Without it, the job replays only the newest 1,000 (see Edge cases).
v2 passes when all four of these hold:
- The control run changes 0 decisions.
- Neither run reports an error or
capped, the flag for a window that held more calls than--max-records. - The blast radius lists each action the candidate added or removed, with its call count; the LexQ console calls each one an effect. It must hold one pair:
MANUAL_REVIEWremoved andAUTO_APPROVEadded, on the same number of calls. A higher limit can only add approvals, so any other effect means v2 changed more than the limit. - The job keeps a sample of the changed calls. Opened one by one, each must be a refund above $50 and up to $75 from an account that passes the other two checks.
A window replay is billed per record, one for each stored call it replays. This check billed 4,000 records, half of them for the control run. Replayed decisions stay in the job's result and never reach the refunds service. Nor are they added to the execution history, the production calls that later replays read: 2026-10-07 replayed again the next day still holds 2,000 calls.
Decision Trace output
Each stored call keeps its decision trace, the record of how each rule fared on that call, and the replay reads its baseline from there. lexq replay get --id <job-id> returns the v2 job's result. Below are its counts and blast radius, with the other fields left out:
{
"totalCount": 2000,
"processedCount": 2000,
"errorCount": 0,
"changedCount": 407,
"capped": false,
"summary": {
"recordsReplayed": 2000,
"recordsChanged": 407,
"determinism": "DETERMINISTIC",
"effectBlastRadius": [
{
"action": {
"type": "SET_FACT",
"parameters": { "targetVar": "refundRoute", "value": "MANUAL_REVIEW" }
},
"direction": "REMOVED",
"recordCount": 407
},
{
"action": {
"type": "SET_FACT",
"parameters": { "targetVar": "refundRoute", "value": "AUTO_APPROVE" }
},
"direction": "ADDED",
"recordCount": 407
}
]
},
"changedSamples": [ ... ]
}
totalCount: calls the job set out to replay, which is the window up to the cap.processedCount: calls replayed so far.changedCount: calls whose actions differ between the stored decision and v2.capped: whether the window held more calls than--max-records.determinism: whether v2 can answer the same input differently on another run. Its rules read nothing but the request's facts.effectBlastRadius: the blast radius, each action that appears or disappears with the number of calls it affects.changedSamples: a sample of the changed calls, each atraceIdwith its effect changes.
A single replay reruns one stored call on the candidate. Run on one changed sample, a $64.99 refund from an 84-day-old account with no recent refunds, it shows why the decision moved.
lexq replay decision --trace-id 6bd369f9-dd46-410c-b5fa-b338b09c8a52 \
--version-id <v2-version-id>
{
"traceId": "6bd369f9-dd46-410c-b5fa-b338b09c8a52",
"decisionChanged": true,
"determinism": "DETERMINISTIC",
"effectChanges": [
{
"action": {
"type": "SET_FACT",
"parameters": { "targetVar": "refundRoute", "value": "MANUAL_REVIEW" }
},
"direction": "REMOVED"
},
{
"action": {
"type": "SET_FACT",
"parameters": { "targetVar": "refundRoute", "value": "AUTO_APPROVE" }
},
"direction": "ADDED"
}
],
"baselineFired": [
{ "ruleName": "Manual review: everything else", ... }
],
"candidateFired": [
{ "ruleName": "Auto-approve: up to $75, established account", ... }
]
}
baselineFired is read from the stored call's decision trace and lists only the rule that won there, the catch-all. The full trace, in the figure below, records the $50 rule as NO_MATCH, because (refundAmountUsd <= 50) was false for 64.99. candidateFired comes from running v2 now. The comparison looks at actions only. A refund of $50 or less that v1's rule approved gets the same AUTO_APPROVE from v2's renamed clone. It does not count as a change.
Edge cases
The default cap is 1,000 calls, and it keeps the newest. Without --max-records, the same day replays only the later 1,000 refunds and reports 184 changed, with capped set. Read as the day's figure, that understates the change by more than half. The cap goes up to 50,000 calls; a window holding more calls than that has to be split into shorter ones.
The window 2026-10-07 to 2026-10-08 spans v2's deploy, so it holds decisions from both versions. Replayed on v2 after 2026-10-08 ends, it reports 407 of 4,000 changed, though v2 is live. All 407 are day-1 decisions that v1 made. With v2 live, that replay is itself the control run, and over a window that spans a deploy it is not 0. A v3 that keeps the $75 limit and changes something else would report those 407 on top of its own change. Windows are whole days in the organization's time zone, so a deploy day with traffic before and after the deploy holds both versions. For the next change after v2, start the window on 2026-10-08, v2's first full day.
Each refund is one plain call, but replay treats three other kinds of stored calls differently. Composite calls, which evaluate several policy groups in one request, are left out of the window. A batch call, which sends many requests at once, stores one row without per-item facts. The replay runs that row on no facts, where only the catch-all fires, and compares its MANUAL_REVIEW with every route in the batch. A call that failed stored its facts but no actions, so any action the candidate takes counts as added. Batch and failed calls can therefore count as changes even in a control run. When a control run is not 0 and no deploy falls in the window, check the window's execution history for failed and batch calls first.
v2 reads only the three facts the refunds service already sends. A candidate that also read a chargeback count would still run in a window replay, though no stored call carries one. The replay treats a condition on the missing fact as not met. The rule does not fire, and the call can count as a change. A single replay of one such call stops with an error instead, which makes it a quick check before the window runs.
The job gives finance no dollar total for the 407 refunds. A changed sample's amount shows only when its trace is opened.
A dispute over one refund is a question about one call, so it takes a single replay like the one in Decision Trace output. Deciding subscription proration with a time-travel audit replays a disputed May charge on July's version to ask whether today's policy would decide it differently.
Production rollout
v2 goes live with one Deploy rather than an A/B ramp. The replay has already run the candidate on every refund of a full day. A split would send identical $60 refunds down different routes for days. The memo on the v2 deploy names both replay jobs and their counts, so the deployment record says why $75 shipped.
Day 1 v1 live, 2,000 refunds stored
control run on v1: 0 of 2,000 changed
candidate run on v2: 407 of 2,000 changed
deploy v2; the memo cites both replay jobs
Day 2 v2 live all day, 2,000 refunds stored
Day 3 replay day 2 on v1: 429 of 2,000 changed
Once v2 has served a full day, the same tool runs in reverse. Replaying that day on v1 compares only decisions v2 made, because no deploy falls inside the window. For 2026-10-08, v1 would have decided 429 of v2's 2,000 refunds differently, with AUTO_APPROVE removed and MANUAL_REVIEW added on 429 calls each. That is 21% against a predicted 20%. Roll back on either of these:
- The blast radius shows any effect other than
AUTO_APPROVEremoved andMANUAL_REVIEWadded. The deployed rules then differ from the replayed ones. v2 was replayed as a draft. Publishing, a separate step before any deploy, locks a version's rules, and until then the draft stays editable. - The changed share climbs well above the 20% the replay predicted. More refunds are arriving just under the new limit.
Neither shows on 2026-10-08, so v2 stays. A rollback would re-point the group at v1 and write its own deployment record. A replay does not time the rules. If the refund endpoint slows after the deploy, per-rule profiling shows whether the two routing rules account for the time.
See how LexQ works for yourself in the playground.
Ready to move decisions out of your deploy pipeline?
Try LexQ free, no credit card required.
Start Free