Portable canonical v3 evidence · decision report

Retain RRF.

The preregistered alternatives improve pooled known-relevant recall, but every alternative loses at least one topic at depth 1,000. That violates the zero-loss promotion rule.

3,953RRF known relevant @1,000
+347primary DUAL pooled gain
20/0/2topic wins / ties / losses
−7worst loss, Topic 31

Why not promote? Primary DUAL loses 7 known-relevant documents on Topic 31 and 3 on Topic 300 at 1,000. RRF500-REINIT-DUAL removes the Topic 31 loss but still loses 3 on Topic 300.

How much graded evidence do we capture?

At depth 20, the authenticated postmortem diagnostic finds 316 / 12,984 known-relevant documents (2.43%) and 2,492 / 84,560 graded-gain units (2.95%). This diagnostic is bound to the pinned qrels and RRF ranking but does not alter the sealed evaluation depths. At depth 1,000, RRF captures 3,953 / 12,984 known-relevant documents (30.45%) and 28,815 / 84,560 graded-gain units (34.08%). The full accepted union raises those ceilings to 38.66% binary recall and 42.38% graded recall.

Known-relevant binary recall gives every judgment at grade 2 or higher equal weight. Graded gain uses 2**grade - 1, so higher-grade evidence contributes more. These are separate questions and neither is precision over unjudged documents.

DepthMetric sourceKnown relevantBinary recallGraded gainGraded recall
20Postmortem diagnostic316 / 12,9842.43%2,492 / 84,5602.95%
100Sealed evaluation1,400 / 12,98410.78%10,532 / 84,56012.46%
250Sealed evaluation2,152 / 12,98416.57%16,116 / 84,56019.06%
500Sealed evaluation2,980 / 12,98422.95%22,024 / 84,56026.05%
1,000Sealed evaluation3,953 / 12,98430.45%28,815 / 84,56034.08%
1,500Sealed evaluation4,526 / 12,98434.86%32,666 / 84,56038.63%
Full unionSealed evaluation5,020 / 12,98438.66%35,840 / 84,56042.38%

Judgment-pool dependent. The denominators cover known judgments only. Unjudged candidates remain unknown, not nonrelevant, and the full-union values are a ceiling within this accepted candidate pool rather than corpus-wide recall.

Why Topics 31 and 300 fail at the cutoff

The pooled gain hides two decisive regressions. Open the native disclosures for the incoming/outgoing label mix, source-family and facet-depth attribution, and the bounded replay. No document text or document identifiers are exposed.

Topic 31 cutoff mechanics: -7 known relevant at 1,000

The candidate moves 322 documents across the cutoff. Outgoing: 322 total · 33 known relevant · 2 judged below 2 · 287 unjudged (unknown). Incoming: 322 total · 26 known relevant · 0 judged below 2 · 296 unjudged (unknown).

Outgoing attribution

Source families: facet only 105; original and facet 18; original only 199.
Best facet-rank buckets: 1 50 101; 101 150 6; 151 200 6; 51 100 10; none 199.
Facet retrieval-stream memberships (actual accepted-union provenance):

  • 31-business-actions (business actions), query businesses practical steps responsible waste handling: 21 memberships · 4 known relevant · 1 judged below 2 · 16 unjudged (unknown) · ranks 1 50 17; 101 150 1; 51 100 3.
  • 31-economies (economies), query proper waste recycling benefits local economies: 22 memberships · 1 known relevant · 0 judged below 2 · 21 unjudged (unknown) · ranks 1 50 15; 101 150 4; 151 200 2; 51 100 1.
  • 31-environment (environment), query e-waste other waste environmental impacts: 10 memberships · 1 known relevant · 0 judged below 2 · 9 unjudged (unknown) · ranks 1 50 4; 101 150 2; 151 200 1; 51 100 3.
  • 31-health (health), query e-waste other waste health impacts: 10 memberships · 1 known relevant · 0 judged below 2 · 9 unjudged (unknown) · ranks 1 50 5; 101 150 1; 51 100 4.
  • 31-improper-management (improper management), query improper waste management risks environmental health: 10 memberships · 0 known relevant · 0 judged below 2 · 10 unjudged (unknown) · ranks 1 50 9; 101 150 1.
  • 31-individual-actions (individual actions), query individuals practical steps responsible waste handling: 25 memberships · 5 known relevant · 1 judged below 2 · 19 unjudged (unknown) · ranks 1 50 14; 101 150 3; 151 200 4; 51 100 4.
  • 31-innovations (innovations), query waste recycling recent innovations: 32 memberships · 0 known relevant · 0 judged below 2 · 32 unjudged (unknown) · ranks 1 50 32.
  • 31-sustainability (sustainability), query proper waste recycling benefits sustainability: 14 memberships · 0 known relevant · 0 judged below 2 · 14 unjudged (unknown) · ranks 1 50 9; 101 150 2; 151 200 2; 51 100 1.

DUAL selection-coverage facets (greedy selection audit, not retrieval-stream provenance): environment 322.

Incoming attribution

Source families: facet only 248; original only 74.
Best facet-rank buckets: 1 50 25; 101 150 84; 151 200 65; 51 100 74; none 74.
Facet retrieval-stream memberships (actual accepted-union provenance):

  • 31-business-actions (business actions), query businesses practical steps responsible waste handling: 23 memberships · 4 known relevant · 0 judged below 2 · 19 unjudged (unknown) · ranks 101 150 11; 151 200 4; 51 100 8.
  • 31-economies (economies), query proper waste recycling benefits local economies: 30 memberships · 1 known relevant · 0 judged below 2 · 29 unjudged (unknown) · ranks 1 50 4; 101 150 12; 151 200 5; 51 100 9.
  • 31-environment (environment), query e-waste other waste environmental impacts: 27 memberships · 3 known relevant · 0 judged below 2 · 24 unjudged (unknown) · ranks 1 50 2; 101 150 10; 151 200 9; 51 100 6.
  • 31-health (health), query e-waste other waste health impacts: 32 memberships · 7 known relevant · 0 judged below 2 · 25 unjudged (unknown) · ranks 1 50 5; 101 150 7; 151 200 11; 51 100 9.
  • 31-improper-management (improper management), query improper waste management risks environmental health: 63 memberships · 1 known relevant · 0 judged below 2 · 62 unjudged (unknown) · ranks 1 50 8; 101 150 17; 151 200 18; 51 100 20.
  • 31-individual-actions (individual actions), query individuals practical steps responsible waste handling: 22 memberships · 2 known relevant · 0 judged below 2 · 20 unjudged (unknown) · ranks 1 50 1; 101 150 3; 151 200 9; 51 100 9.
  • 31-innovations (innovations), query waste recycling recent innovations: 16 memberships · 2 known relevant · 0 judged below 2 · 14 unjudged (unknown) · ranks 1 50 1; 101 150 8; 151 200 3; 51 100 4.
  • 31-sustainability (sustainability), query proper waste recycling benefits sustainability: 39 memberships · 0 known relevant · 0 judged below 2 · 39 unjudged (unknown) · ranks 1 50 4; 101 150 16; 151 200 10; 51 100 9.

DUAL selection-coverage facets (greedy selection audit, not retrieval-stream provenance): economies 1; environment 320; innovations 1.

Topic 300 cutoff mechanics: -3 known relevant at 1,000

The candidate moves 281 documents across the cutoff. Outgoing: 281 total · 16 known relevant · 37 judged below 2 · 228 unjudged (unknown). Incoming: 281 total · 13 known relevant · 11 judged below 2 · 257 unjudged (unknown).

Outgoing attribution

Source families: facet only 102; original only 179.
Best facet-rank buckets: 1 50 34; 101 150 7; 51 100 61; none 179.
Facet retrieval-stream memberships (actual accepted-union provenance):

  • 300-antarctica (antarctica), query climate change actions help Antarctica: 26 memberships · 0 known relevant · 1 judged below 2 · 25 unjudged (unknown) · ranks 1 50 10; 101 150 3; 51 100 13.
  • 300-economic-costs (economic costs), query global warming economic costs addressing compared dealing impacts: 32 memberships · 0 known relevant · 0 judged below 2 · 32 unjudged (unknown) · ranks 1 50 11; 101 150 3; 51 100 18.
  • 300-global-measures (global measures), query global measures address global warming climate change: 30 memberships · 3 known relevant · 27 judged below 2 · 0 unjudged (unknown) · ranks 1 50 7; 51 100 23.
  • 300-strategies (strategies), query global warming climate change effective strategies prevent reduce: 14 memberships · 5 known relevant · 8 judged below 2 · 1 unjudged (unknown) · ranks 1 50 6; 101 150 1; 51 100 7.

DUAL selection-coverage facets (greedy selection audit, not retrieval-stream provenance): strategies 281.

Incoming attribution

Source families: facet only 172; original only 109.
Best facet-rank buckets: 101 150 94; 151 200 78; none 109.
Facet retrieval-stream memberships (actual accepted-union provenance):

  • 300-antarctica (antarctica), query climate change actions help Antarctica: 35 memberships · 0 known relevant · 7 judged below 2 · 28 unjudged (unknown) · ranks 101 150 22; 151 200 13.
  • 300-economic-costs (economic costs), query global warming economic costs addressing compared dealing impacts: 32 memberships · 2 known relevant · 0 judged below 2 · 30 unjudged (unknown) · ranks 101 150 17; 151 200 15.
  • 300-global-measures (global measures), query global measures address global warming climate change: 51 memberships · 0 known relevant · 0 judged below 2 · 51 unjudged (unknown) · ranks 101 150 28; 151 200 23.
  • 300-strategies (strategies), query global warming climate change effective strategies prevent reduce: 54 memberships · 1 known relevant · 0 judged below 2 · 53 unjudged (unknown) · ranks 101 150 27; 151 200 27.

DUAL selection-coverage facets (greedy selection audit, not retrieval-stream provenance): antarctica 2; economic costs 2; global measures 1; strategies 276.

Topic 300 facet-tail replay: cap 100 recovers +2 at 1,000

The offline replay protects the first 100 RRF results, keeps original-stream candidates eligible, and defers facet-only candidates whose best facet rank is deeper than cap 100. Outgoing: 183 total · 9 known relevant · 28 judged below 2 · 146 unjudged (unknown). Incoming: 183 total · 11 known relevant · 4 judged below 2 · 168 unjudged (unknown).

Facet-tail known-relevant yield: 1-50 36.18% · 101-150 6.00% · 151-200 8.00% · 51-100 29.29%. The +2 result is retrospective recovery evidence, not a promotion result.

What this result means—and what it does not

This is a retrospective full-development stress test over all 22 development topics using already known judgments. It measures retrieval of known-relevant documents and is useful for diagnosing headroom and regressions. It is not evidence of generalization to new topics, unseen judgments, or production traffic. The downstream RAG answer generation is out of scope; no claim is made about answer accuracy, faithfulness, citation quality, or user utility.

The v1/v2 ranking/evaluation rejected label is intentional: only portable rankings_v3 and evaluation_v3 are admissible here. Their ranking bytes and scientific conclusion are unchanged from corrected v2; v3 removes checkout-local paths from sealed identity and enforces the judged-rate promotion guard.

Exact preregistered selection ladder

The baseline appears first for orientation; alternatives follow the sealed ladder exactly. Aggregate improvements are insufficient when a zero-loss guard fails.

ArmKnown rel. @1kΔ vs RRFPooled recallW/T/LLoss topicsDecision
RRF3,953+030.4%0/22/0noneretain baseline
RRF100-STATIC-DUAL4,300+34733.1%20/0/231, 300reject
RRF500-REINIT-DUAL4,290+33733.0%21/0/1300reject
RRF100-REINIT-DUAL4,300+34733.1%20/0/231, 300reject
RRF100-REINIT-DUAL-NR4,312+35933.2%20/0/231, 300reject
RRF100-STATIC-DUAL-NR4,312+35933.2%20/0/231, 300reject

Recall and retained facet evidence by depth

Recall rises as the cutoff expands, and DUAL variants retain more facet-only evidence. The aggregate curves explain the attraction of the alternatives; the per-topic guard below explains the decision.

Recall and retained facet evidence by depthBinary recall rises with depth for RRF and the two strongest DUAL variants. Alternatives improve aggregate recall but still cause topic regressions.0.00.10.20.31002505001,0001,500Full union
RRFPrimary DUALRRF500 DUAL
Pooled binary recall over the 12,984 known-relevant documents. Depth 100 is protected and identical across arms.

Judged-rate caveat: RRF judged rate falls from 97.8% at 100 to 28.2% at 1,000. Unjudged documents are not negatives, so deep precision and yield are conservative and judgment-pool dependent.

Known-relevant yield falls with facet depth

The first 50 results of each deduplicated facet stream carry the highest pooled known-relevant yield. Later buckets still add evidence but at lower density, supporting depth discipline rather than blanket expansion.

Known-relevant yield falls with facet depthPooled known-relevant yield by facet rank bucket, calculated across all twenty-two topics.16.7%1-5012.6%51-10010.1%101-1508.6%151-200
Pooled known-relevant count divided by pooled unique candidates in each facet-rank bucket across 148 facet streams.

Per-topic deltas expose the promotion blockers

Every topic is shown. Count deltas are alternative minus RRF; negative values are regressions. The final column gives primary DUAL deltas at depths 250 / 500 / 1,500.

TopicKnown rel.RRF @1kRRF recallPrimary Δ @1kRRF500 Δ @1kNR Δ @1kPrimary Δ 250/500/1500
Topic 1469518626.8%+3+7+5+5 / +5 / +3
Topic 3192527329.5%-7+5-5-27 / -18 / +6
Topic 3782715719.0%+20+23+23+12 / +14 / +23
Topic 5851623645.7%+34+35+37+16 / +16 / +34
Topic 7290126529.4%+7+21+9-6 / -14 / +24
Topic 8470118526.4%+7+14+3-10 / -6 / +5
Topic 14442112930.6%+27+21+27+32 / +28 / +34
Topic 16168725437.0%+19+15+20+13 / +15 / +10
Topic 2004947615.4%+11+8+11+11 / +12 / +9
Topic 2131737543.4%+29+29+29+29 / +36 / +12
Topic 21958010017.2%+10+5+10-5 / +2 / +10
Topic 22465321032.2%+16+12+15+9 / +17 / +15
Topic 22555525545.9%+22+17+18-10 / -6 / +6
Topic 23378726934.2%+21+19+22+4 / +24 / +0
Topic 27333010030.3%+4+4+8+2 / +14 / +0
Topic 30063522635.6%-3-3-2+20 / +17 / +8
Topic 4072967625.7%+8+9+8+9 / +11 / +1
Topic 47752514527.6%+28+18+31+0 / +13 / +15
Topic 49960935658.5%+25+25+24-15 / -15 / +2
Topic 5154146615.9%+21+14+22+15 / +17 / +20
Topic 70758613222.5%+27+21+28+8 / +18 / +14
Topic 89767418227.0%+18+18+16+15 / +18 / +24

Representative diagnostics

Topic 31: primary DUAL -7 at 1,000

RRF retrieves 273 of 925 known-relevant documents at depth 1,000 (29.51%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are -27, -18, -7, and +6.

Topic 37: primary DUAL +20 at 1,000

RRF retrieves 157 of 827 known-relevant documents at depth 1,000 (18.98%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +12, +14, +20, and +23.

Topic 58: primary DUAL +34 at 1,000

RRF retrieves 236 of 516 known-relevant documents at depth 1,000 (45.74%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +16, +16, +34, and +34.

Topic 144: primary DUAL +27 at 1,000

RRF retrieves 129 of 421 known-relevant documents at depth 1,000 (30.64%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +32, +28, +27, and +34.

Topic 213: primary DUAL +29 at 1,000

RRF retrieves 75 of 173 known-relevant documents at depth 1,000 (43.35%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +29, +36, +29, and +12.

Topic 225: primary DUAL +22 at 1,000

RRF retrieves 255 of 555 known-relevant documents at depth 1,000 (45.95%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are -10, -6, +22, and +6.

Topic 233: primary DUAL +21 at 1,000

RRF retrieves 269 of 787 known-relevant documents at depth 1,000 (34.18%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +4, +24, +21, and +0.

Topic 300: primary DUAL -3 at 1,000

RRF retrieves 226 of 635 known-relevant documents at depth 1,000 (35.59%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +20, +17, -3, and +8.

Topic 477: primary DUAL +28 at 1,000

RRF retrieves 145 of 525 known-relevant documents at depth 1,000 (27.62%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +0, +13, +28, and +15.

Topic 499: primary DUAL +25 at 1,000

RRF retrieves 356 of 609 known-relevant documents at depth 1,000 (58.46%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are -15, -15, +25, and +2.

Topic 515: primary DUAL +21 at 1,000

RRF retrieves 66 of 414 known-relevant documents at depth 1,000 (15.94%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +15, +17, +21, and +20.

Topic 707: primary DUAL +27 at 1,000

RRF retrieves 132 of 586 known-relevant documents at depth 1,000 (22.53%). Primary DUAL deltas at depths 250, 500, 1,000, and 1,500 are +8, +18, +27, and +14.

Method and evidence boundary

1 · Plan22 narratives → 148 tethered facet queries.
2 · RetrieveTop 200 per facet; 45,144-document union.
3 · ScoreLocal MiniLM narrative/facet features.
4 · Freeze blindSix complete arms, qrels unopened.
5 · EvaluatePinned development qrels opened only after v3 freeze verification.
6 · DecideApply exact ladder and all promotion guards.

The portable ranking freeze has 22 topics, 6 complete arms, 270,864 ranking rows, and 206,030 audit rows. The evaluation independently binds that freeze before opening the pinned qrels. The paired exact sign-flip test for primary DUAL estimates +3.394 percentage points mean topic recall (95% bootstrap CI +2.118 to +5.052; raw p=0.000002623; Holm-adjusted p=0.000006676). Statistical significance cannot override the preregistered topic-loss rule.

Costs and execution accounting

148live facet retrieval requests
771,725local forward pairs
2727.437slocal scoring wall time
$0 recordedhosted / paid inference

Original-query requests: 0. Retrieval failures/retries: 0/0. Shared-score cache reuses: 68,079. Completed score windows: 848,183. Peak device/host memory: 489,715,200 / 12,323,328,000 bytes. Ranking and evaluation added zero retrieval, inference, model-load, hosted, paid, or network calls.

Limitations, decision, and next step

Recommendation: retain RRF, run a bounded recovery lane for Topics 31/300-like failures, and start the RAG lane from a deeper pool with Source-diverse evidence selection. Do not pass the protected top 20 through unchanged and call it facet-aware RAG.

Authenticated provenance

Ranking v3 root: 6f4c35899f90c1d60324caf24bf8834d3482c5e8b9e785f8316ab0eec55fc305
Evaluation v3 root: e49ce3f7f0cfeedf1863670ae4cfe5781d40163449a172a5822d8eb11eeffe84
Pinned qrels SHA-256: 42bf933ae06eb22213312b22e3f2bc39f3dcc2d54e87ebcd8125e9528ddfcc37
Planning root: bc1351cca8aa05dd0979a342c8dd5f72395f2b7a9207ae20668aba05f1ab85a2
Retrieval root: f2191c295f600d0243f4dcd2b1dc23a05b9fa8a7d9a4d6258983a0d3c9433aeb
Scoring root: 4f67107770b1cd35598adc75beeeafe205ba984f201cedf4b505c4d20515e990

This sanitized report contains aggregate metrics and hashes only: no credentials, raw qrels, raw documents, document identifiers, request logs, or local filesystem paths.

Back to top