从控制到动态平衡:超级智能对齐的博弈论与热力学重构
Current research on AI alignment mainly follows paths such as reinforcement learning from human feedback, scalable oversight, and interpretability. Its implicit premise is that there exists a centralized designer capable of defining clear objectives and exercising effective control over the system. However, when AI sys...