Skip to content

Author

Simiao Ren

We have 6 of 16 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Review Aug 2026

A Corpus of Real Scam- and Spam-Call Conversations from an Active Voice-Agent Honeypot

A dataset of real scam-call conversations collected by an active voice-agent honeypot, describing the collection system, record structure, and technical validation of the corpus's realism and label quality, including that the agent is recognized as non-human in only about 5% of engaged calls.

Ethan Traister, D. Ng, Si-Yu Zhang et al. · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control

AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what that means for a relying party -- an insurer, a lender, an auditor -- whose evidence is a filed PDF. AgentForge-Bench measures how...

Si-Miao Ren, Ankit Raj, Tommy Duong et al. · 0 citations
#natural language process... Preprint Sep 2026

Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model

Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds. Jev-style typed decisions promise exactly that: declared options go in, one calibrated probability per option comes out of a single forward pass, with no generated text. We test an open implementation of this read...

Si-Miao Ren, Kidus Zewde, Xing-Yu Shen et al. · 6 citations · ⚡1
#artificial intelligence Review Sep 2026

ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation

An image tool can change its underlying generator while retaining its public name, making version attribution from online posts ambiguous. We study this problem after the ChatGPT Images 2.5 launch. Our frozen collection contains 3,478 images from 2,440 posts across 8 sources. Recorded posting times fall within the firs...

Dennis Ng, Xing-Yu Shen, Ankit Raj et al. · 1 citation
Preprint Aug 2026

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

This work presents CallScreenBench, which reports five automated call-and-note measure groups motivated by owner endorsement, and reports quality measures and guardedness channels separately so that a single pass/fail score does not hide their trade-offs.

Jia-Qi Gan, Hao Tang, Jamey Z. Liang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.