LLM agents take real actions, including executing code, modifying files, calling services, and delegating tasks, driven by context sources: user requests, tool results, documents, shell outputs, Skill and MCP instructions, memory. Unlike traditional systems, where capability is predefined, the least-privilege capabilit...
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising findi...
Yu-Zhu Mao, Lei Yu, Zi-Ning Zhu et al.· 0 citations
AgentPProf is a profiler that aggregates agent trajectories into pprof-compatible profiles, enabling flame graph visualization and analysis and introduces recursive operation segmentation, which recursively splits trajectories at task boundaries.
Yu-Sheng Zheng, Chaokun Chang, Yuan-Man Mao et al.· 0 citations
KernelScript is presented, a DSL that types maps, program handles, and execution domains in one source, then compiles to standard C through the original toolchain, and rejects cross-boundary bugs at compile time that standard C/libbpf still builds and loads.
NetArtifactBench is introduced, which tests whether AI agents can repair inconsistent records derived from public network-system artifacts while preserving claims that remain supported, and argues that artifact integrity should become a first-class design and evaluation requirement for AI agents operating on network sy...
Tianzhu Zhang, Wei-Chen Tao, Chang-Gang Zheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.