Skip to content

Author

Yangfan Song

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching

Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware methods that produce a different compressed prefix per query, mechanically invalidating t...

Yangfan Song · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.