MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning
This work builds an agent harness that separates reasoning from perception and improves by evolving context rather than optimizing weights, and proposes a gradient-free, reward-gated Heuristic Skill Distillation loop that mines the agent's own low-scoring traces and keeps a candidate skill only when it raises a validation reward, yielding reusable retrieval skills, notably directed re-look.