Skip to content
#explainable ai Open access

Harbor: A framework for evaluating and optimizing agents and models in container environments

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

What's Changed Fix FX installation and background persistence by @DarlingtonDeveloper in https://github.com/harbor-framework/harbor/pull/2829 Migrate agents to structured capability declarations by @alexgshaw in https://github.com/harbor-framework/harbor/pull/2834 fix(rewardkit): trajectory criteria default to /logs/agent/trajectory.json by @no-hup in https://github.com/harbor-framework/harbor/pull/2840 fix(islo): slug ephemeral gateway profile names by @tomerezer in https://github.com/harbor-framework/harbor/pull/2844 fix(vercel): auto-load .env.local so onboarding auth flow doesn't loop by @erulkey in https://github.com/harbor-framework/harbor/pull/2845 fix(qwen-code): preserve reasoning and subagent trajectories by @DragonnZhang in https://github.com/harbor-framework/harbor/pull/2364 Fix/hyperbrowser remote image builds by @Dingway98 in https://github.com/harbor-framework/harbor/pull/2833 Docs trajectory loading by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2859 Docs handoff by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2860 docs regrade by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2864 Add swe interact example task by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2865 Google Sans Code logo by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2861 Docs simulated user by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2866 docs: remote rollouts guide for hosted Harbor by @scvance in https://github.com/harbor-framework/harbor/pull/2867 refactor(hub)!: rename the secret's provider field to --label by @scvance in https://github.com/harbor-framework/harbor/pull/2862 docs: add usage stats by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2884 docs: job results viewer by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2889 fix: extend mini-swe LiteLLM timeout to 3600s by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2886 fix(hosted): make credential selection explicit for hosted launches by @scvance in https://github.com/harbor-framework/harbor/pull/2857 feat: add Podman as a ContainerRuntime behind DockerEnvironment by @xichen1997 in https://github.com/harbor-framework/harbor/pull/2872 docs: results handoff by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2890 docs: add custom agents guide by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2892 docs: add pre-integrated agents by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2891 Docs acp agents by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2894 docs: use Geist and Google Sans Code on the Mintlify site by @alexgshaw in https://github.com/harbor-framework/harbor/pull/2828 fix doc by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2897 docs: add Harbor tasks overview by @alexgshaw in https://github.com/harbor-framework/harbor/pull/2898 docs: add ATIF trajectory by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2899 docs: "uv run harbor" -> "harbor" by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2901 docs: -p -> -t by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2900 fix(viewer): prefer canonical reward in job results by @ayushnangia in https://github.com/harbor-framework/harbor/pull/2877 feat(tensorlake sandbox): dynamic network policy + docker-in-docker support by @cooleel in https://github.com/harbor-framework/harbor/pull/2817 Simplify nested RewardKit groups by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/2836 Use kebab-case RewardKit aggregations by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/2906 fix: classify new Anthropic refusals pattern by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2922 fix(env): do not classify TOKENS as secrets by @scvance in https://github.com/harbor-framework/harbor/pull/2925 Fix Windows cross-drive artifact paths by @MarcoRossignoli in https://github.com/harbor-framework/harbor/pull/2878 feat(atif): support audio content parts (ATIF-v1.8) by @miguelrc-scale in https://github.com/harbor-framework/harbor/pull/2605 feat(computer-1): support Gemini 3.7 Flash computer use by @erikqu in https://github.com/harbor-framework/harbor/pull/2936 docs: add environment variable by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2926 docs: add artifact collection by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2937 docs: add custom sandbox by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2935 docs: add CLI, config, Python to preinstalled agents by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2934 fix(verifier): reject non-finite (NaN/Infinity) reward values at ingestion by @no-hup in https://github.com/harbor-framework/harbor/pull/2888 Add fx dev-channel agent by @fazxes in https://github.com/harbor-framework/harbor/pull/2932 Add configurable session ID headers by @shariqm-modal in https://github.com/harbor-framework/harbor/pull/2723 Detect likely Modal sandbox OOM exits by @scvance in https://github.com/harbor-framework/harbor/pull/2955 docs: add existing plugins guide by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2957 docs: pre-integrated sandboxes by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2933 Validate wheel extras before publish by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2960 docs: add release policy by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2971 docs: explain separate verifier environment variables by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2970 Improve mintlify navigation by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2969 docs: add custom plugins by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2959 docs: add versioning by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2966 fix(rewardkit): don't double-register a reweighted zero-parameter criterion by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/2945 Fix ACP bootstrap on older task images by @DarlingtonDeveloper in https://github.com/harbor-framework/harbor/pull/2973 fix(skills): Handle multiple git remote urls from git remote-ls in skill resuolution by @cdxker in https://github.com/harbor-framework/harbor/pull/2958 fix: remove tensorlake from cloud extra by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2980 docs: reduce ATIF wording by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2902 Default harbor auth login to the manual flow in SSH sessions by @devin-ai-integration[bot] in https://github.com/harbor-framework/harbor/pull/2981 fix(openhands-sdk): forward reasoning effort to LiteLLM by @sravell in https://github.com/harbor-framework/harbor/pull/2927 fix: restore tensorlake in cloud extra by @cooleel in https://github.com/harbor-framework/harbor/pull/2991 docs: illustrate network policy phases by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2990 docs: document agent capabilities by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2989 docs: clarify changelog in CONTRIBUTING.md by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2988 Docs curated changelog by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2982 docs: update stable release policy by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2984 fix(docker): stream stdout without a line-length limit by @somu84 in https://github.com/harbor-framework/harbor/pull/2985 feat(acp): support task-provided local agents by @andy-zhou in https://github.com/harbor-framework/harbor/pull/2992 opencode: opt out of OAuth for remote MCP servers by @psbang in https://github.com/harbor-framework/harbor/pull/2918 antigravity-sdk: forward GOOGLE_GEMINI_BASE_URL to the runner by @psbang in https://github.com/harbor-framework/harbor/pull/2917 Bound the Antigravity model discovery command timeout by @alexgshaw in https://github.com/harbor-framework/harbor/pull/2999 Pin mini-swe-agent installer Python by @RyanMarten in https://github.com/harbor-framework/harbor/pull/3008 Add fx agent judge to RewardKit by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/2848 fix(rewardkit): decide tests/ layout from registered criteria by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/2944 feat(rewardkit): signed weighted aggregation and validated judge TOMLs by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/2879 docs: fill asp sandbox page with rfc 0003 content by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3024 feat: add kata environment for KVM microVM isolation by @xichen1997 in https://github.com/harbor-framework/harbor/pull/2998 small fix on asp rfc. by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3026 docs: drop stale pinned version from README citation block by @StevenDillmann in https://github.com/harbor-framework/harbor/pull/3027 Run HF sandbox commands with Bash by @evalstate in https://github.com/harbor-framework/harbor/pull/3016 feat(claude-code): add thinking_display kwarg (--thinking-display) by @hrdkbhatnagar in https://github.com/harbor-framework/harbor/pull/3030 Support composing job config files by @walkerhughes in https://github.com/harbor-framework/harbor/pull/3022 docs: add Hosted Harbor Mintlify section by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3051 Add Muse Code installed agent integration by @think-step-by-step in https://github.com/harbor-framework/harbor/pull/3043 Improve agent error classification by @scvance in https://github.com/harbor-framework/harbor/pull/3055 fix(acp): make managed Python readable by agent by @andy-zhou in https://github.com/harbor-framework/harbor/pull/3015 Declare and validate built-in

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#artificial intelligence Conference Open access Jun 2018

The Key Concepts of Ethics of Artificial Intelligence

It is suggested that the focus on finding keywords is the first step in guiding and providing direction for future research in the AI ethics field.

Ville Vakkuri, P. Abrahamsson · 39 citations · ⚡2

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.