Skip to content

Author

Francesco Marchetti

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Sep 2026

EuLLM — Open-source sovereign LLM platform

EuLLM is an open-source platform for creating, distributing, and running sovereign EU-compliant Large Language Models, designed for verticalization across domains, languages, and brands while ensuring AI Act compliance. The platform consists of three components: Engine — a Rust-based inference runtime built on top of llama.cpp, exposing OpenAI-compatible and Ollama-compatible APIs on the same port, with TurboQuant KV cache compression for up to 4× context length on consumer GPUs and a continuous batching scheduler for parallel decode of multiple concurrent requests; Forge — a verticalization pipeline written in Python that compresses 14B-30B foundation models down to specialized 7B domain experts via structural pruning, knowledge distillation, quantization, and identity LoRA fine-tuning; Hub — an EU-hosted model registry providing AI Act compliance cards, provenance documentation, and verifiable model lineage tracking, distributed via European cloud infrastructure to ensure data sovereignty. Use cases include sovereign EU LLM deployment for regulated industries (legal, healthcare, finance), domain-specific verticalized models, and EU AI Act compliant inference infrastructure.

Francesco Marchetti · 0 citations
#large language models Open access Sep 2026

EuLLM — Open-source sovereign LLM platform

EuLLM is an open-source platform for creating, distributing, and running sovereign EU-compliant Large Language Models, designed for verticalization across domains, languages, and brands while ensuring AI Act compliance. The platform consists of three components: Engine — a Rust-based inference runtime built on top of llama.cpp, exposing OpenAI-compatible and Ollama-compatible APIs on the same port, with TurboQuant KV cache compression for up to 4× context length on consumer GPUs and a continuous batching scheduler for parallel decode of multiple concurrent requests; Forge — a verticalization pipeline written in Python that compresses 14B-30B foundation models down to specialized 7B domain experts via structural pruning, knowledge distillation, quantization, and identity LoRA fine-tuning; Hub — an EU-hosted model registry providing AI Act compliance cards, provenance documentation, and verifiable model lineage tracking, distributed via European cloud infrastructure to ensure data sovereignty. Use cases include sovereign EU LLM deployment for regulated industries (legal, healthcare, finance), domain-specific verticalized models, and EU AI Act compliant inference infrastructure.

Francesco Marchetti · 0 citations
#large language models Open access Sep 2026

EuLLM — Open-source sovereign LLM platform

EuLLM is an open-source platform for creating, distributing, and running sovereign EU-compliant Large Language Models, designed for verticalization across domains, languages, and brands while ensuring AI Act compliance. The platform consists of three components: Engine — a Rust-based inference runtime built on top of llama.cpp, exposing OpenAI-compatible and Ollama-compatible APIs on the same port, with TurboQuant KV cache compression for up to 4× context length on consumer GPUs and a continuous batching scheduler for parallel decode of multiple concurrent requests; Forge — a verticalization pipeline written in Python that compresses 14B-30B foundation models down to specialized 7B domain experts via structural pruning, knowledge distillation, quantization, and identity LoRA fine-tuning; Hub — an EU-hosted model registry providing AI Act compliance cards, provenance documentation, and verifiable model lineage tracking, distributed via European cloud infrastructure to ensure data sovereignty. Use cases include sovereign EU LLM deployment for regulated industries (legal, healthcare, finance), domain-specific verticalized models, and EU AI Act compliant inference infrastructure.

Francesco Marchetti · 0 citations
#large language models Open access Sep 2026

EuLLM — Open-source sovereign LLM platform

EuLLM is an open-source platform for creating, distributing, and running sovereign EU-compliant Large Language Models, designed for verticalization across domains, languages, and brands while ensuring AI Act compliance. The platform consists of three components: Engine — a Rust-based inference runtime built on top of llama.cpp, exposing OpenAI-compatible and Ollama-compatible APIs on the same port, with TurboQuant KV cache compression for up to 4× context length on consumer GPUs and a continuous batching scheduler for parallel decode of multiple concurrent requests; Forge — a verticalization pipeline written in Python that compresses 14B-30B foundation models down to specialized 7B domain experts via structural pruning, knowledge distillation, quantization, and identity LoRA fine-tuning; Hub — an EU-hosted model registry providing AI Act compliance cards, provenance documentation, and verifiable model lineage tracking, distributed via European cloud infrastructure to ensure data sovereignty. Use cases include sovereign EU LLM deployment for regulated industries (legal, healthcare, finance), domain-specific verticalized models, and EU AI Act compliant inference infrastructure.

Francesco Marchetti · 0 citations
#large language models Open access Sep 2026

EuLLM — Open-source sovereign LLM platform

EuLLM is an open-source platform for creating, distributing, and running sovereign EU-compliant Large Language Models, designed for verticalization across domains, languages, and brands while ensuring AI Act compliance. The platform consists of three components: Engine — a Rust-based inference runtime built on top of llama.cpp, exposing OpenAI-compatible and Ollama-compatible APIs on the same port, with TurboQuant KV cache compression for up to 4× context length on consumer GPUs and a continuous batching scheduler for parallel decode of multiple concurrent requests; Forge — a verticalization pipeline written in Python that compresses 14B-30B foundation models down to specialized 7B domain experts via structural pruning, knowledge distillation, quantization, and identity LoRA fine-tuning; Hub — an EU-hosted model registry providing AI Act compliance cards, provenance documentation, and verifiable model lineage tracking, distributed via European cloud infrastructure to ensure data sovereignty. Use cases include sovereign EU LLM deployment for regulated industries (legal, healthcare, finance), domain-specific verticalized models, and EU AI Act compliant inference infrastructure.

Francesco Marchetti · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.