← Jiacheng Xu
EMNLP 2026. (Oral, ~2.6% of submissions)

ERPO

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

Jiacheng Xu, Feng Chen, Xiuneng Xu, and Bo An

Nanyang Technological University, Singapore

Test-time learning from program behavior, without oracle outputs. ERPO uses noisy consensus conservatively to improve both single-sample accuracy and the coverage of correct solutions across multiple samples.

ERPO pipeline: generate output-free probe inputs, execute candidate programs to compute behavioral consensus rewards, then mask the higher-ranked half and apply an entropy-regularized update to the lower-ranked candidates.
Figure 2 from the paper: probe generation, behavioral scoring, and conservative policy optimization. View full-size figure.

Abstract

Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal.

To make TTRL applicable to code generation, we propose probe-driven TTRL, which constructs output-free probe inputs from the problem statement, executes candidate programs on these probes, and defines a Probe Consensus Reward (PCR) from the resulting behavioral agreement. PCR provides a behavioral training signal for open-vocabulary programs, but it is not a fully reliable verifier and remains susceptible to reward hacking through spurious consensus.

We therefore introduce Entropy-Regularized Rank-Masked Policy Optimization (ERPO), which converts low PCR into conservative negative updates through rank masking and controls policy drift with an entropy ceiling. On coding benchmarks, ERPO substantially improves pass@1 and pass@k in both in-domain adaptation and zero-shot transfer.

Research topics: test-time reinforcement learning, code generation, execution-based feedback, behavioral consensus, reward hacking, and inference-time scaling.

Probe-driven TTRL

Agreement is a signal, not a correctness certificate.

Two correct programs may look different, while two incorrect programs may share the same bug. ERPO compares execution behavior and adapts the policy without treating a popular output as ground truth.

Unlike training-free answer selection, ERPO updates model parameters on unlabeled target problems. Hidden tests are reserved for evaluation, not adaptation.

Step 1 · Probe inputs

Generate shared execution inputs

Generate diverse inputs from each problem statement without predicting their correct outputs. The experiments use 10 generated inputs per problem, combined with statement example inputs and deduplicated. Probes are computed once and reused during adaptation.

Step 2 · Probe Consensus Reward

Score behavioral agreement

Execute candidate programs on the same probes. PCR is the fraction of probes on which a candidate's output belongs to the majority set. Tied majority outputs all receive credit; execution failures receive none and are excluded from majority counting.

Step 3 · ERPO update

Mask high-ranked candidates and cap entropy growth

Normalize PCR within each candidate group and mask the top-ranked half. Retain the signed advantages of the lower-ranked half, which are usually negative, to primarily suppress low-consensus programs. A fixed entropy ceiling penalizes excessive entropy growth, while leaving entropy below the ceiling unpenalized.

Qwen3-4B and Qwen3-8B · non-thinking mode

Gains on the adaptation benchmark and unseen target benchmarks.

Adaptation uses unlabeled LiveCodeBench v6 problems. The resulting checkpoint is then evaluated on CodeContests, CodeForces, and TACO without updates on those transfer benchmarks.

Selected rows from Table 1. All scores are percentages; transfer mean averages CodeContests, CodeForces, and TACO.

Swipe or scroll the table to see all results.

Base and ERPO comparison (%)
Model / methodLCB
pass@1
LCB
pass@16
Transfer mean
pass@1
Transfer mean
pass@16
Qwen3-4B · Base26.034.413.825.7
Qwen3-4B · ERPO36.746.325.439.3
Qwen3-8B · Base25.634.614.627.6
Qwen3-8B · ERPO35.150.922.938.8

For Qwen3-4B on LiveCodeBench, ERPO improves pass@1 by 10.7 percentage points and pass@16 by 11.9 points over the base model. Table 1 in the paper also reports GRPO with public-test rewards, GRPO with PCR, and negative-sample reinforcement with PCR.

How to read pass@k: it measures the probability of obtaining at least one correct solution among k samples. It is an oracle coverage metric, not the accuracy of a practical selector. These main results use 16 validation samples per problem; the separate 256-sample scaling experiment is reported in Table 3.

Learning with noisy feedback still needs care.

PCR measures agreement, not correctness: shared bugs and inadequate probes can produce high scores, while a correct minority solution can receive a low score. ERPO mitigates these risks through conservative updates; it does not make PCR a trusted verifier.

The experiments cover competitive-programming-style tasks with Qwen3-4B and Qwen3-8B. Project-level and industrial-scale code generation were not evaluated. Adaptation also has a compute cost: the paper reports approximately 12 hours per full training run on eight 80 GB NVIDIA A100 GPUs. Improved single-sample generation should not be read as cost-free adaptation.

Related work: Test Cases Scaling (TCS) learns sound and adversarial test generation using reference solutions. ERPO instead studies test-time policy adaptation using output-free probes and noisy behavioral consensus. Both investigate execution-based feedback for code language models, under different supervision settings.

Accepted to EMNLP 2026 as an Oral presentation (~2.6% of submissions). The public manuscript was first submitted to arXiv on 8 September 2026. The PDF is an unchanged copy of arXiv v1; the method image is rendered from its original figure. Paper and figure by Jiacheng Xu, Feng Chen, Xiuneng Xu, and Bo An, available under CC BY 4.0.

Jiacheng Xu, Feng Chen, Xiuneng Xu, and Bo An. “Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation.” 2026. arXiv:2609.09135.

The citation below identifies the public arXiv version.

@article{xu2026erpo,
  title   = {Entropy-Regularized Rank-Masked Policy Optimization for
             Test-Time Reinforcement Learning in Code Generation},
  author  = {Jiacheng Xu and Feng Chen and Xiuneng Xu and Bo An},
  journal = {arXiv preprint arXiv:2609.09135},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.09135}
}