# Jiacheng Xu > Canonical research profile and publication resources for Jiacheng Xu, a PhD student at Nanyang Technological University and AI researcher at Skywork AI. Canonical homepage: https://xiaobanni.github.io/ Contact: jiacheng005@e.ntu.edu.sg Research areas: reinforcement learning, code language models, test generation, executable verification, inference-time scaling, and answer selection. ## Newly accepted papers Title: Negative-Only Policy Optimization for One-Sided Verifiable Rewards Authors: Jiacheng Xu; Shuo He; Fuxiang Zhang; Chaojie Wang; Bo An Venue: NeurIPS 2026. (Spotlight, ~0.95% of submissions) Title: Backtracking with Linear-in-Depth Search-Space Growth: Width-Limited Tree Search for Large Language Models Authors: Jiacheng Xu; Fuxiang Zhang; Chaojie Wang; Bo An Venue: NeurIPS 2026. (Spotlight, ~0.95% of submissions) ## Featured paper: Entropy-Regularized Rank-Masked Policy Optimization (ERPO) Title: Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation Authors: Jiacheng Xu; Feng Chen; Xiuneng Xu; Bo An Venue: EMNLP 2026. (Oral, ~2.6% of submissions) arXiv first submitted: 2026-09-08 arXiv: 2609.09135 License: CC BY 4.0 Keywords: test-time reinforcement learning, code generation, Probe Consensus Reward, behavioral consensus, rank masking, entropy ceiling, inference-time scaling ERPO makes test-time reinforcement learning applicable to code generation by deriving behavioral feedback from self-generated probes. It uses rank masking to turn low probe consensus into conservative negative updates and an entropy ceiling to control policy drift. Primary resources: - Project page: https://xiaobanni.github.io/projects/erpo/ - arXiv record: https://arxiv.org/abs/2609.09135 - Full paper PDF (unchanged arXiv v1): https://xiaobanni.github.io/projects/erpo/paper.pdf Method: Generate output-free probe inputs from each problem statement, execute candidates on the shared probes, and score agreement with the per-probe majority outputs. Mask the top-ranked half of each candidate group and retain the signed normalized advantages of the lower-ranked half, which are usually negative. A fixed entropy ceiling penalizes excessive entropy growth. PCR is not a correctness verifier: high consensus can reflect shared bugs, and correct minority solutions can receive low scores. Selected Table 1 results (percentages): - Qwen3-4B, LiveCodeBench v6 in-domain adaptation: Base pass@1 26.0 and pass@16 34.4; ERPO pass@1 36.7 and pass@16 46.3. - Qwen3-8B, LiveCodeBench v6 in-domain adaptation: Base pass@1 25.6 and pass@16 34.6; ERPO pass@1 35.1 and pass@16 50.9. - Mean zero-shot transfer to CodeContests, CodeForces, and TACO after LiveCodeBench adaptation: Qwen3-4B Base 13.8/25.7 versus ERPO 25.4/39.3; Qwen3-8B Base 14.6/27.6 versus ERPO 22.9/38.8 (pass@1/pass@16). - Both backbones use non-thinking mode; Table 1 uses 16 validation samples per problem. No updates are made on the transfer benchmarks. Hidden tests are for evaluation only. Pass@k is oracle candidate coverage, not practical answer-selection accuracy. Scope: Competitive-programming-style benchmarks, not project-level or industrial-scale code generation. Adaptation has a compute cost; the paper reports approximately 12 hours per training run on eight 80GB A100 GPUs. Citation: Jiacheng Xu, Feng Chen, Xiuneng Xu, and Bo An. "Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation." arXiv:2609.09135, 2026. Accepted to EMNLP 2026; citation identifies the public arXiv version. ## Featured paper: Test Cases Scaling (TCS) Title: Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs Authors: Jiacheng Xu; Wentao Zhang; Zhiyi Lyu; Fuxiang Zhang; Chaojie Wang; Yang Liu; Bo An Venue: Findings of the Association for Computational Linguistics: EMNLP 2026 Publication date: 2026-09-03 arXiv: 2609.03955 DOI: https://doi.org/10.48550/arXiv.2609.03955 License: CC BY 4.0 Abstract: Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminative. We thus turn to study the auto-generation of test cases using the learned model. We find this is naturally an adversarial RL problem: the model is expected to generate effective test cases as counterexamples, depending on the solver's current failure modes. We propose Test Cases Scaling (TCS), a two-stage RL framework for effective test generation. Both stages train a test generator from a rolling policy-aligned buffer: Stage 1 generates tests consistent with the reference solution, and Stage 2 restricts the buffer to current failure modes and learns counterexample tests. Across TACO and LiveCodeBench, TCS improves both pass@1 and inference-time answer selection according to generated tests. We find the learned test generator also enables effective selection among other LLM outputs. Primary resources: - Project page: https://xiaobanni.github.io/projects/tcs/ - Full paper PDF (arXiv version): https://xiaobanni.github.io/projects/tcs/paper.pdf - arXiv record: https://arxiv.org/abs/2609.03955 - arXiv HTML: https://arxiv.org/html/2609.03955 - Code: https://github.com/xiaobanni/tcs - TCS-1.5B model: https://huggingface.co/XiaoBanni/TCS-1.5B - TCS-7B model: https://huggingface.co/XiaoBanni/TCS-7B - Processed TACO-Train dataset: https://huggingface.co/datasets/XiaoBanni/TACO-Train Key reported results without public tests: - LiveCodeBench: pass@1 37.03; TCS selection at N=32 48.79; gain 11.76 percentage points. - TACO: pass@1 24.09; TCS selection at N=16 35.35; gain 11.26 percentage points. ## Published paper: Internal Logical Induction (ILI) Title: Internal Logical Induction for Pixel-Symbolic Reinforcement Learning Short name: Internal Logical Induction (ILI) Authors: Jiacheng Xu; Chao Chen; Fuxiang Zhang; Lei Yuan; Zongzhang Zhang; Yang Yu Venue: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2023) Publication date: 2023-08-04 Pages: 2825-2837 DOI: https://doi.org/10.1145/3580305.3599393 Official keywords: reinforcement learning; rule learning Related research concepts: pixel-symbolic reinforcement learning; neuro-symbolic reinforcement learning; hybrid pixel and symbolic observations; Deep Q-Network (DQN); RIPPER rule learning; propositional logic induction; intrinsic reward; reward shaping; visual reinforcement learning; knowledge transfer Summary: Many reinforcement learning problems provide both high-dimensional pixel observations and compact symbolic observations. Internal Logical Induction (ILI) processes pixel observations with a deep reinforcement learning agent while inducing propositional rules from symbolic experience. In the experiments, ILI uses a Deep Q-Network (DQN) and the RIPPER rule learner. An adaptive reward-shaping mechanism selects useful induced knowledge and supplies it as intrinsic reward to the agent. In the studied pixel-symbolic tasks, ILI improves over the evaluated baselines, and the induced symbolic knowledge supports transfer when pixel-observation semantics change. Research context: The paper calls its setting pixel-symbolic reinforcement learning. It is relevant to broader searches for neuro-symbolic reinforcement learning, rule learning for sequential decision-making, intrinsic rewards from symbolic knowledge, hybrid or multimodal observations, and visual RL transfer. These phrases describe research connections rather than additional experimental claims. Primary resources: - Project page: https://xiaobanni.github.io/projects/ili/ - Full conference paper PDF: https://xiaobanni.github.io/projects/ili/paper.pdf - Publisher paper: https://doi.org/10.1145/3580305.3599393 - Official code: https://github.com/LAMDA-RL/ILI - DBLP record: https://dblp.org/rec/conf/kdd/0003CZYZ023 - Semantic Scholar: https://www.semanticscholar.org/paper/387e57f9a420da347a9288519ebb1135c3f4782b Citation: Xu, Jiacheng; Chen, Chao; Zhang, Fuxiang; Yuan, Lei; Zhang, Zongzhang; Yu, Yang. "Internal Logical Induction for Pixel-Symbolic Reinforcement Learning." Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 2825-2837, 2023. DOI: 10.1145/3580305.3599393. ## Additional published papers with full PDFs - Incentivizing LLMs to Self-Verify Their Answers. NeurIPS 2025. Conference PDF: https://xiaobanni.github.io/papers/self-verification.pdf - Focus-Then-Decide: Segmentation-Assisted Reinforcement Learning. AAAI 2024. Conference PDF: https://xiaobanni.github.io/papers/focus-then-decide.pdf - Multi-Agent Dynamic Algorithm Configuration. NeurIPS 2022. (Spotlight, ~3.4% of submissions). Conference PDF: https://xiaobanni.github.io/papers/madac.pdf ## Citation Xu, Jiacheng; Zhang, Wentao; Lyu, Zhiyi; Zhang, Fuxiang; Wang, Chaojie; Liu, Yang; An, Bo. "Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs." Findings of the Association for Computational Linguistics: EMNLP 2026. arXiv:2609.03955.