State2State: Environment-Derived Mid-Training for LLM Agents
State2State is proposed, an environment-derived mid-training method that converts explored environment states into training objectives, challenging agents to reach a specified target state by deriving tasks from environment exploration and verifying success through rule-based state matching.