Author

Guangtao Nie

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

CEComBench: Benchmarking Large Language Models' performance on Chinese E-commerce tasks

We introduce CEComBench (Chinese E-Commerce Benchmark), a rigorously curated evaluation framework comprising 12140 annotated samples spanning 36 distinct tasks, sourced from JD.com, a leading Chinese e-commerce platform. Crucially, our data collection, task generation, and evaluation pipeline eschew LLM involvement to mitigate potential biases and ensure consistency. Instead, domain experts and human annotators are systematically engaged to uphold benchmark quality and neutrality. CEComBench serves as a substantial contribution to the existing landscape of E-Commerce benchmarks due to its large scale, high quality, and real world data sources, as well as its objective generation and evaluation. With rigorous experiments of trending LLMs such as GPT4, Claude, Qwen series, and DeepSeek series, we reveal several findings that challenge common scaling assumptions. We uncover a fundamental gap between generation fluency and reasoning ability, identify a pronounced ''inverse scaling effect'' where larger models can underperform in domain-specific reasoning, and pinpoint systemic bottlenecks across all SOTA models, such as a ''Structure Barrier'' in complex data extraction, exposing fundamental limitations of current architectures. The benchmark is now publicly available at https://huggingface.co/datasets/jdopensource/CEComBench.

Guangtao Nie, Huimu Wang, Gewei Lu et al. · 0 citations