Internal Logical Induction for Pixel-Symbolic Reinforcement Learning
Deep reinforcement learning and propositional rule induction for environments with mixed pixel and symbolic observations.
Affiliations at publication: Nanjing University · Polixir Technologies
ILI learns propositional rules from symbolic experience and converts selected rules into intrinsic rewards that guide a deep RL agent learning from pixels.
Why combine pixels and symbols?
Reinforcement learning agents are often designed around a single observation type. Pixel observations contain rich perceptual detail but are high-dimensional, while symbolic observations provide compact, semantically meaningful descriptions. In many environments both are available, yet treating them identically does not make full use of their complementary strengths.
Internal Logical Induction (ILI) separates these roles. A deep reinforcement learning component learns from pixel observations, and a rule-learning component induces propositional knowledge from symbolic interaction data. ILI then uses an adaptive reward-shaping mechanism to select useful induced knowledge and provide it to the RL agent as intrinsic reward.
In the studied pixel-symbolic reinforcement learning tasks, ILI outperforms the evaluated baselines. The learned propositional knowledge also helps transfer when the meaning of the pixel observations changes while the underlying symbolic structure remains useful.
Official paper keywords: Reinforcement Learning; Rule Learning.
Perception
A deep RL policy learns decision-making directly from high-dimensional pixel observations.
Induction
A rule learner extracts compact propositional knowledge from symbolic experience.
Transfer
Symbolic knowledge can remain useful when the semantics of the pixel observations change.
Internal Logical Induction
Learn rules internally.
Use them as rewards.
The neural and symbolic components are connected through reward shaping rather than by forcing both observation types into the same representation.
In the experiments, ILI instantiates the components with a Deep Q-Network (DQN) and the RIPPER rule-learning algorithm; the framework itself is designed to accommodate other RL and rule learners.
-
1
Observe pixels and symbols
The environment supplies a pixel state for the deep RL agent and a symbolic state for knowledge induction.
-
2
Induce propositional rules
Symbolic interaction data is used to learn logical knowledge from successful experience.
-
3
Select valuable knowledge
An adaptive mechanism identifies induced knowledge that is useful for the agent's current learning process.
-
4
Shape the RL reward
Selected logical knowledge becomes intrinsic reward, guiding the pixel-based deep RL policy.
Terminology bridge
How this paper relates to nearby research areas
The paper names its setting pixel-symbolic reinforcement learning. Its method is also relevant when searching the broader literature below; these terms describe connections, not additional experimental claims.
- Neuro-symbolic reinforcement learning
- ILI combines a neural deep RL policy with symbolic propositional rule induction in one learning framework.
- Hybrid or multimodal observations
- The task setting provides both raw visual information and compact symbolic features with explicit semantics.
- Intrinsic rewards and reward shaping
- Induced logical knowledge is not merely reported: selected knowledge is converted into intrinsic reward for policy learning.
- Visual RL generalization and transfer
- The paper studies whether learned symbolic knowledge can aid policy transfer when pixel-input semantics change.
- Rule learning for sequential decision-making
- Propositional rules summarize useful symbolic experience and provide a compact source of guidance for reinforcement learning.
Jiacheng Xu, Chao Chen, Fuxiang Zhang, Lei Yuan, Zongzhang Zhang, and Yang Yu. “Internal Logical Induction for Pixel-Symbolic Reinforcement Learning.” KDD 2023, pages 2825–2837. DOI: 10.1145/3580305.3599393.
@inproceedings{xu2023internal,
title = {Internal Logical Induction for Pixel-Symbolic
Reinforcement Learning},
author = {Jiacheng Xu and Chao Chen and Fuxiang Zhang and
Lei Yuan and Zongzhang Zhang and Yang Yu},
booktitle = {Proceedings of the 29th ACM SIGKDD Conference on
Knowledge Discovery and Data Mining},
pages = {2825--2837},
year = {2023},
doi = {10.1145/3580305.3599393}
}