Skip to content
#explainable ai Open access

Harbor: A framework for evaluating and optimizing agents and models in container environments

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management

Abstract

What's Changed Publish local launch tasks through the existing publisher by @scvance in https://github.com/harbor-framework/harbor/pull/3184 fix(environments): name the actual build-context dir in the missing-definition error by @hellno in https://github.com/harbor-framework/harbor/pull/2854 fix(codex): preserve web-search IDs in ATIF by @ayushnangia in https://github.com/harbor-framework/harbor/pull/2972 fix(computer-1): send prompt_cache_key so OpenAI runs hit the cache by @neverSettles in https://github.com/harbor-framework/harbor/pull/3190 fix(viewer): show all rewards in jobs list hover tooltip by @thealchen in https://github.com/harbor-framework/harbor/pull/3201 fix(rewardkit): json_key_equals must not treat a missing key as null by @Rome-1 in https://github.com/harbor-framework/harbor/pull/3182 fix(computer-1): execute native Anthropic, Bedrock, and Gemini batches by @erikqu in https://github.com/harbor-framework/harbor/pull/3204 feat: stream trajectories and sandbox editor over SSH by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3171 docs: explain streaming logs and sandbox files by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3172 fix(multi-step): don't leak prior step's /tests and /logs/verifier into the next agent phase by @EachSheep in https://github.com/harbor-framework/harbor/pull/1961 fix: avoid OverflowError in environment_content_hash for files >4GB by @axiom-of-choice in https://github.com/harbor-framework/harbor/pull/3011 Add --use-static-ip flag for hosted launches by @scvance in https://github.com/harbor-framework/harbor/pull/3213 Add Code Puppy agent (based on Pydantic AI) by @mpfaffenberger in https://github.com/harbor-framework/harbor/pull/3199 Warn when separate verifiers exclude configured mounts by @ayushnangia in https://github.com/harbor-framework/harbor/pull/3193 Add first-class Muse Code installed agent support by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/3166 Group RewardKit criteria by Python file by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/2947 feat: declare and validate agent skills and MCP capabilities by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3220 refactor: share SSH setup helpers by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3228 chore: limit reviewer mentions to rewardkit by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3221 feat(tensorlake): enforce compose network policy at the sandbox level by @cooleel in https://github.com/harbor-framework/harbor/pull/3225 docs: document agent MCP and skills capabilities by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3248 docs: replace stream screenshot with demo video by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3250 docs: add v0.23.0 changelog highlights by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3251 docs: correct prompt_template_path parameter name by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3252 docs: fix leaderboard and network policy links by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3253 feat: support Modal streaming by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3229 docs: use Mintlify trees for task directory examples by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3242 docs: add upgrade and nightly installation instructions by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3243 docs: highlight optional environment directory for prebuilt images by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3245 docs: link MCP and skills compatibility to agent capabilities by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3246 refactor: move Windows capability into _validate_agent_capabilities by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3255 fix(viewer): keep trial tab bar in view when switching tabs by @thealchen in https://github.com/harbor-framework/harbor/pull/3233 Port Rewardkit docs to Mintlify by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/3234 docs: expand sidebar navigation groups by default by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3256 docs: render sidebar sections as fixed headings by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3265 refactor(cli): move add and remove under dataset with hidden aliases by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3261 docs: redirect legacy documentation to Mintlify by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3259 Bump harbor-rewardkit to 0.2.1 by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/3270 docs: correct hosted launch and custom-agent guidance by @scvance in https://github.com/harbor-framework/harbor/pull/3241 docs: add docs MCP sidebar link by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3266 fix: read goose thinking blocks from the thinking field by @alran in https://github.com/harbor-framework/harbor/pull/3274 fix: preserve executable status in task archives by @scvance in https://github.com/harbor-framework/harbor/pull/3282 fix(viewer): prevent sticky tabs from covering stream controls by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3284 Add --launch support to trial and job regrade by @scvance in https://github.com/harbor-framework/harbor/pull/3237 fix: install only missing agent system dependencies by @scvance in https://github.com/harbor-framework/harbor/pull/3277 fix: record Hub dataset downloads by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3291 fix (stream): write stream.json after agent setup by @cooleel in https://github.com/harbor-framework/harbor/pull/3257 fix(antigravity-cli): load MCP servers declared in task.toml on agy >= 1.1.27 by @xa8zz in https://github.com/harbor-framework/harbor/pull/3095 Update Rewardkit docs for per-file scoring by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/3269 Reject non-finite RewardKit scores by @ayushnangia in https://github.com/harbor-framework/harbor/pull/3218 fix(rewardkit): don't score a missing CSV column as an empty match by @no-hup in https://github.com/harbor-framework/harbor/pull/3272 Fail the trial when a GKE pod is gone instead of returning exit code 1 by @mgeorgaklis in https://github.com/harbor-framework/harbor/pull/3215 Honor task build timeout for Modal Compose startup by @scvance in https://github.com/harbor-framework/harbor/pull/3305 fix: close Daytona client before event loop shutdown by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3302 docs: move News into the docs and migrate existing announcements by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3286 docs: add June–September 2026 news announcements by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3287 docs: ask fork PRs to allow maintainer edits by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3309 feat: --effort by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3304 fix: classify launch failures from harness events only by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3216 feat: add a common disable_web_search agent option by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3308 feat(antigravity-cli): add error markers, options by @yaoshengzhe in https://github.com/harbor-framework/harbor/pull/3307 fix(beam): update sandbox integration by @luke-lombardi in https://github.com/harbor-framework/harbor/pull/3306 fix(rewardkit): clean up judge processes on cancellation by @ayushnangia in https://github.com/harbor-framework/harbor/pull/3300 Improve viewer evaluation tables, filters, and trial navigation by @alexgshaw in https://github.com/harbor-framework/harbor/pull/3311 feat(rewardkit): add JEV judge and rubric criteria by @benediktstroebl in https://github.com/harbor-framework/harbor/pull/3325 fix(hermes): setup-timeout floor, oneshot export, drop the 90-turn cap and compression override by @teknium1 in https://github.com/harbor-framework/harbor/pull/3321 feat(cwsandbox): merge the wandb environment into cwsandbox by @matthoare117-wandb in https://github.com/harbor-framework/harbor/pull/3085 feat(tensorlake): add streaming support for tensorlake sandboxes by @cooleel in https://github.com/harbor-framework/harbor/pull/3281 fix(pi): support max thinking level by @scvance in https://github.com/harbor-framework/harbor/pull/3345 docs(runta): restore Runta in Mintlify documentation by @SASUKE40 in https://github.com/harbor-framework/harbor/pull/3342 feat(agents): add codex and opencode as simulated-user targets by @joebaumann in https://github.com/harbor-framework/harbor/pull/3343 feat(pi-agent): add MCP, turn limits, and ATIF trajectories by @stackviolator in https://github.com/harbor-framework/harbor/pull/3317 fix(pi): preserve thinking levels for custom endpoints by @scvance in https://github.com/harbor-framework/harbor/pull/3352 docs: note streaming sandbox use for debugging by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3347 docs: add missing streaming sandbox providers by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3348 docs: remove legacy docs by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3354 docs: document Pi capabilities in Mintlify by @kobe0938 in https://github.com/harbor-framework/harbor/pull/3367 fix(docker): add SELinux label support for bind mounts by @not-stbenjam in https://github.com/harbor-framework/harbor/pull/1991 feat: harbor run diff by @kobe0938 in https://github.com/harbor-framework/harbor/pull/2430 feat(trial): expose resolved timeout properties by @0oshowero0 in https://github.com/harbor-framework/harbor/pull/3388 fix(viewer): show the RewardKit judge model by @ayushnangia in https://github.com/harb

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#artificial intelligence Conference Open access Jun 2018

The Key Concepts of Ethics of Artificial Intelligence

It is suggested that the focus on finding keywords is the first step in guiding and providing direction for future research in the AI ethics field.

Ville Vakkuri, P. Abrahamsson · 39 citations · ⚡2

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.