Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training
Tan, B., Wang, J., Hou, D., Jiang, L., Wu, Z., Shen, Y., Lin, F., Yamada, K., & Koike, A. (2026). Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training. arXiv preprint arXiv:2608.08224.
