Skip to content
Book Open access

Code Generation from Regression Trees for Microsecond-Scale Decisions in Operating Systems

Sep 2026 · Proceedings of the 14th Workshop on Programming Languages and Operating Systems · pp. 86-93 · 0 citations · 11 references

TL;DR

This work examines the suitability of regression trees for fast (nano- to microsecond-scale) runtime decisions within operating systems, and is able to reduce inference latency by up to an order of magnitude compared to conventional approaches.

Abstract

In today's world of heterogeneous server hardware, deciding on suitable task and data placements is a far from trivial undertaking. Depending on system load and application behaviour, some compute and memory assignments improve performance, while others impair it. Yet, these increasingly complex decisions must be made quickly: spending 5 ms just to decide on a placement that reduces latency by merely 2 ms is a net loss. So far, developers have resorted to hand-crafted heuristics for this task. While these are founded in real-world experience, they are hard to reason about, hard to adjust, and rarely free from bugs or inefficiencies. In our opinion, a transition towards model-guided placement decisions is due - i.e., using machine learning models to predict compute/data placement from system status and application requirements. To this end, we examine the suitability of regression trees for fast (nano- to microsecond-scale) runtime decisions within operating systems. We examine three types of trees / forests with three different code generation methods, and show that, depending on tree complexity and data type, median latency is 23 ns to 79 &mgr;s per placement decision. By using C++ template expansion to transform entire trees into equivalent machine code, we are able to reduce inference latency by up to an order of magnitude compared to conventional approaches, while (for 8-bit integer data) also halving the code size. At the same time, trees are compact and can easily be generated from, e.g., Python-based machine learning models.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

PerfReasoning is introduced, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code, exposing the gap between plausible architectural reasoning and reliable performance-model construction.

Da Zhao, K. Sankaralingam, Christos Kozyrakis et al. · 0 citations
#natural language process... Preprint Sep 2026

StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams

Language models increasingly make real-time decisions in applications that apply the latest answer until a newer one arrives. A late answer can prolong an outdated decision, such as a call recorder still running while a customer reads out card details, an error offline accuracy misses. We make three contributions. Firs...

Jhen-Ke Lin, Chung-Chun Wang · 0 citations
Open access Sep 2026

Architectural Foundations for Latency-Aware Scalability in AI-Enhanced Enterprise Systems

This paper examines two production-documented architectures that fold detection, decision, verification, and staged deployment into one closed loop, one built for continuous security enforcement and the other for continuous performance optimization, and asks whether the integration amounts to a genuine architectural ad...

Daniil Sergeevich Martynov · 0 citations
Book Open access Aug 2026

Critical Path Guided Decision Making with CALLIGATOR

Modern web applications are typically implemented by many independent microservices. Seeing into such a distributed system and predicting how component changes propagate across the system is essential for informed decision-making. Unfortunately, existing tools provide only partial support for such analysis in microserv...

Meghna Pancholi, Lee Baugh, Olaf Schnapauff et al. · 0 citations
Preprint Aug 2026

Aneto: Predicting System Performance by Exploiting Cross-Workload Regularity

Aneto is a mechanistic-empirical regression model that estimates the performance-latency sensitivity of any new workload from a single run, enabling first-order CPI prediction under any memory configuration.

Raúl Taranco, Rene Mueller, Michael Giardino · 0 citations
Preprint Sep 2026

Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization

Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g., GPUs), the CPU increasingly acts as an orchestrator, spending cycles in driver calls, data movement, and synchronization rather than in application code. Se...

Vlad-Petru Nitu, Harsh Songara, Konstantinos Sgouras et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.