How Do OpenHands Coding Agents Use Tools? Best of 2024 in Agents with Prof. Graham Neubig (#1 on SWE-Bench Full, OpenHands/AllHands)

13.5K views
•
December 25, 2024
by
Latent Space
YouTube video player
How Do OpenHands Coding Agents Use Tools? Best of 2024 in Agents with Prof. Graham Neubig (#1 on SWE-Bench Full, OpenHands/AllHands)

TL;DR

OpenHands coding agents build and improve software by browsing the web, editing files, running Bash commands and Jupyter cells, and writing programs that use APIs and existing libraries. Prof. Graham Neubig demonstrates three practical tasks: analyzing SWE-Bench results, creating an email script, and extending a monitoring application. Read on to see how concrete prompts, programmatic tool use, and inspectable actions make these agents effective.

Transcript

okay uh hi everyone so um I was given the task of talking about agents in 2024 and this is an impossible task because there are so many agents so many agents in 2024 so this is going to be strongly covered by like my personal experience and what I think is interesting and important but um I think it's an important topic so let let's go ahead so um ... Read More

Key Insights

  • Coding agents can support data analysis, new software creation, and iterative improvement of existing applications. Neubig says he uses them five to ten times a day, showing that their practical role extends beyond generating isolated code snippets to completing varied software-development tasks.
  • Concrete prompts improve agent performance. The demonstrations specify outputs, data sources, interfaces, and implementation details, including labeled scatter plots, CSV input, Jinja2 templates, API documentation, pie charts, and totals across a monitored period, giving each agent a clear operational target.
  • Programmatic tool use can reduce repeated model calls. Instead of asking a model to make roughly 30 granular API calls, OpenHands lets the agent write and execute Python code that invokes exposed APIs together, observes errors, and revises the program when necessary.
  • OpenHands relies on a small core tool set. Its agent can execute Bash programs and Jupyter notebook cells, browse and overwrite file sections, perform global search and replacement, and browse the web through actions such as scrolling, clicking, and entering text.
  • Existing software libraries expand an agent's effective capabilities. Access to Python packages for HTTP requests, PDF-to-text conversion, and other tasks allows a coding agent to perform data visualization and related work without requiring a separately designed tool for every possible operation.
  • GitHub integration enables agents to work within active development workflows. An agent can clone repositories, use the GitHub API, inspect issue comments, check GitHub Actions, respond when tagged, and attempt fixes such as repairing failing tests on a pull request.
  • Human-agent interfaces must balance visibility with information overload. OpenHands presents an English description for each action, hides potentially large Bash commands or notebook contents by default, and lets users open those details when closer inspection becomes useful.
  • Coding agents can complete useful work while a user handles another activity. During the presentation, an OpenHands agent analyzed the SWE-Bench repository and generated labeled plots while Neubig continued speaking, illustrating asynchronous assistance within a broader working session.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What software-development tasks can OpenHands coding agents handle?

They can perform data analysis, create new software, and improve existing applications. Graham Neubig demonstrates these capabilities by plotting SWE-Bench score changes, building a CSV-driven email script from API documentation, and adding pie charts and total counts to repository-monitoring software.

Q: What tools does an OpenHands coding agent use?

OpenHands gives agents capabilities for executing Bash programs and Jupyter notebook cells, editing files, performing global search and replacement, and browsing the web. Its browsing actions include scrolling, clicking, and entering text, while its file tools can expose and overwrite selected content.

Q: Why do concrete prompts improve coding-agent performance?

Concrete prompts define the desired output, inputs, constraints, and information sources for the agent. Neubig’s examples specify details such as labeled scatter plots, CSV input, Jinja2 templates, API documentation, pie charts, and totals across the monitored period.

Q: How does OpenHands reduce repeated model-directed tool calls?

OpenHands lets an agent write Python programs that call exposed APIs and libraries together. Instead of directing the model through roughly 30 separate API calls, the agent can execute one program, observe errors, revise the code, and run it again.

Q: How do existing software libraries expand a coding agent’s capabilities?

Agents can use libraries already available to programmers instead of requiring a custom tool for every task. Neubig mentions Python packages for HTTP requests and PDF-to-text conversion, while the demonstrations also show agents using executable code for data visualization and API interaction.

Q: How does OpenHands let users inspect an agent’s work?

OpenHands describes each agent action in English, such as stating that it will create an email-sending script. Lengthy Bash commands and Jupyter notebook contents remain collapsed by default, but users can open them when they need to inspect the details.

Q: How can OpenHands agents work within GitHub workflows?

The OpenHands GitHub plugin can respond when users tag the agent in issues or pull requests. The agent can clone repositories, inspect issue comments and GitHub Actions through the GitHub API, and attempt tasks such as resolving an issue or fixing failing pull-request tests.

Q: What did the live SWE-Bench analysis demonstrate?

An OpenHands agent analyzed the SWE-Bench repository and created scatter plots while Neubig continued presenting. It labeled the top three systems for the requested datasets, including SWE-Bench normal, verified, and light, demonstrating that an agent can complete useful data-analysis work asynchronously.

Summary & Key Takeaways

  • Graham Neubig demonstrates three everyday coding-agent tasks: analyzing SWE-Bench score changes, creating an email-sending script from API documentation, and extending an existing monitoring application. The data-analysis agent finishes during the presentation, producing scatter plots and labeling leading systems across several SWE-Bench datasets while Neubig continues speaking to the audience.

  • OpenHands gives its agent a compact set of capabilities for executing Bash commands and Jupyter cells, editing files, performing global search and replacement, and browsing the web. The agent can then write Python programs that call APIs and libraries, allowing one program execution to replace a long sequence of separate model-directed tool calls.

  • The human-agent interface aims to reveal enough information without overwhelming users. OpenHands describes each action in English, collapses lengthy commands by default, and permits deeper inspection of Bash activity and Jupyter notebooks. It also meets developers in existing workflows through chat, GitHub issue tagging, pull-request assistance, and a remote headless runtime.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Latent Space 📚