LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts'intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs'ability to generate executable commands for real-world cy...
Peng-Fei Li, Naufal Suryanto, Si-Cheng Zhang et al.· 0 citations
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware ev...
Peng-Fei Li, Naufal Suryanto, Si-Cheng Zhang et al.· 0 citations
Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we introduce LingT2I, a benchmark covering 10 widely u...
Si-Cheng Zhang, Zhong-Hao Yan, Bin-Zhu Xie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.