Workplace sensing studies combine long-running behaviour traces with self-reports, yet the tools that collect those data often sit apart from the interface that returns results. We present TrustmeWatcher, the application built for the TRUST-ME project to connect this work. TrustmeWatcher reuses ActivityWatch's OS-level...
Cheng-Yu Yu, Leonor Costa, Zoja Anžur et al.· 0 citations
Synthetic semantic-preserving transformations that are rule-based and invertible are used as a probe of alignment generalization and suggest that semantic-preserving distribution shifts can expose a recurring gap in how utility and alignment generalize in current LLMs.
Mo-Han Li, Cheng-Yu Yu, Francesco Sovrano et al.· 0 citations
As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate such risks by checking proposed actions before execution, but many r...
Wenhao Lin, Cheng-Yu Yu, Xingwei Lin et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.