Test-time learning from program behavior, without oracle outputs. ERPO uses noisy consensus conservatively to improve both single-sample accuracy and the coverage of correct solutions across multiple samples.
Figure 2 from the paper: probe generation, behavioral scoring, and conservative policy optimization. View full-size figure.
Paper overview
Abstract
Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal.
To make TTRL applicable to code generation, we propose probe-driven TTRL, which constructs output-free probe inputs from the problem statement, executes candidate programs on these probes, and defines a Probe Consensus Reward (PCR) from the resulting behavioral agreement. PCR provides a behavioral training signal for open-vocabulary programs, but it is not a fully reliable verifier and remains susceptible to reward hacking through spurious consensus.
We therefore introduce Entropy-Regularized Rank-Masked Policy Optimization (ERPO), which converts low PCR into conservative negative updates through rank masking and controls policy drift with an entropy ceiling. On coding benchmarks, ERPO substantially improves pass@1 and pass@k in both in-domain adaptation and zero-shot transfer.
Research topics: test-time reinforcement learning, code generation, execution-based feedback, behavioral consensus, reward hacking, and inference-time scaling.
01 Method
Probe-driven TTRL
Agreement is a signal, not a correctness certificate.
Two correct programs may look different, while two incorrect programs may share the same bug. ERPO compares execution behavior and adapts the policy without treating a popular output as ground truth.
Unlike training-free answer selection, ERPO updates model parameters on unlabeled target problems. Hidden tests are reserved for evaluation, not adaptation.
Step 1 · Probe inputs
Generate shared execution inputs
Generate diverse inputs from each problem statement without predicting their correct outputs. The experiments use 10 generated inputs per problem, combined with statement example inputs and deduplicated. Probes are computed once and reused during adaptation.
Step 2 · Probe Consensus Reward
Score behavioral agreement
Execute candidate programs on the same probes. PCR is the fraction of probes on which a candidate's output belongs to the majority set. Tied majority outputs all receive credit; execution failures receive none and are excluded from majority counting.
Step 3 · ERPO update
Mask high-ranked candidates and cap entropy growth
Normalize PCR within each candidate group and mask the top-ranked half. Retain the signed advantages of the lower-ranked half, which are usually negative, to primarily suppress low-consensus programs. A fixed entropy ceiling penalizes excessive entropy growth, while leaving entropy below the ceiling unpenalized.
02 Results
Qwen3-4B and Qwen3-8B · non-thinking mode
Gains on the adaptation benchmark and unseen target benchmarks.
Adaptation uses unlabeled LiveCodeBench v6 problems. The resulting checkpoint is then evaluated on CodeContests, CodeForces, and TACO without updates on those transfer benchmarks.
Selected rows from Table 1. All scores are percentages; transfer mean averages CodeContests, CodeForces, and TACO.
Swipe or scroll the table to see all results.
Base and ERPO comparison (%)
Model / method
LCB pass@1
LCB pass@16
Transfer mean pass@1
Transfer mean pass@16
Qwen3-4B · Base
26.0
34.4
13.8
25.7
Qwen3-4B · ERPO
36.7
46.3
25.4
39.3
Qwen3-8B · Base
25.6
34.6
14.6
27.6
Qwen3-8B · ERPO
35.1
50.9
22.9
38.8
For Qwen3-4B on LiveCodeBench, ERPO improves pass@1 by 10.7 percentage points and pass@16 by 11.9 points over the base model. Table 1 in the paper also reports GRPO with public-test rewards, GRPO with PCR, and negative-sample reinforcement with PCR.
How to read pass@k: it measures the probability of obtaining at least one correct solution among k samples. It is an oracle coverage metric, not the accuracy of a practical selector. These main results use 16 validation samples per problem; the separate 256-sample scaling experiment is reported in Table 3.
03 Scope and limitations
Learning with noisy feedback still needs care.
PCR measures agreement, not correctness: shared bugs and inadequate probes can produce high scores, while a correct minority solution can receive a low score. ERPO mitigates these risks through conservative updates; it does not make PCR a trusted verifier.
The experiments cover competitive-programming-style tasks with Qwen3-4B and Qwen3-8B. Project-level and industrial-scale code generation were not evaluated. Adaptation also has a compute cost: the paper reports approximately 12 hours per full training run on eight 80 GB NVIDIA A100 GPUs. Improved single-sample generation should not be read as cost-free adaptation.
Related work:Test Cases Scaling (TCS) learns sound and adversarial test generation using reference solutions. ERPO instead studies test-time policy adaptation using output-free probes and noisy behavioral consensus. Both investigate execution-based feedback for code language models, under different supervision settings.
Accepted to EMNLP 2026 as an Oral presentation (~2.6% of submissions). The public manuscript was first submitted to arXiv on 8 September 2026. The PDF is an unchanged copy of arXiv v1; the method image is rendered from its original figure. Paper and figure by Jiacheng Xu, Feng Chen, Xiuneng Xu, and Bo An, available under CC BY 4.0.
05 Citation
Jiacheng Xu, Feng Chen, Xiuneng Xu, and Bo An. “Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation.” 2026. arXiv:2609.09135.
The citation below identifies the public arXiv version.
@article{xu2026erpo,
title = {Entropy-Regularized Rank-Masked Policy Optimization for
Test-Time Reinforcement Learning in Code Generation},
author = {Jiacheng Xu and Feng Chen and Xiuneng Xu and Bo An},
journal = {arXiv preprint arXiv:2609.09135},
year = {2026},
url = {https://arxiv.org/abs/2609.09135}
}