Speaker
Abstract
Computer Use Agents represent a paradigm shift in software interaction: AI models trained to operate interfaces visually, mimicking human interaction rather than relying on technical APIs. This presentation explores the underlying mechanics of these agents and demonstrates how to orchestrate them programmatically to execute common frontend testing tasks.
The session illustrates how agents — using Claude 3.5 Sonnet, Opus 4.6 and the Chrome-MCP — navigate complex graphical user interfaces, identify creative workarounds for missing functionalities, and document their findings. However, there are significant "visual reasoning gaps" where current models struggle, such as interpreting angular displays like gauges or complex spatial networks like street maps, which can quickly exceed context limits. Attendees will gain a realistic assessment of the performance and reliability of Computer Use Agents in modern automation workflows.
Slides from this presentation can be found here.
Topics
QCon London 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Emerging Trends in the Frontend and Mobile Hosted by Ian Thomas Software Engineer @Meta, QCon London & San Francisco Co-Chair, International SpeakerFrom the same track
Wednesday 18 March
10:35 Mountbatten (6th Fl.) Session AI Tools That Enable the Next 1B Developers Ivan Zarea Director of Platform Engineering @Netlify The developer population is growing by orders of magnitude. The new AI tooling is turning domain experts and operators into builders. Unfortunately, the tools and the platforms they are building on weren't designed for them. 11:45 Mountbatten (6th Fl.) Session GUI Agents Computer Use Agents: The Frontier of Vision-Based Automation Stefan Dirnstorfer CTO @Thetaris GmbH, Architect of ThetaML & Thetaris’ Testing Application, 42 Years into Software Computer Use Agents represent a paradigm shift in software interaction: AI models trained to operate interfaces visually, mimicking human interaction rather than relying on technical APIs. 13:35 Windsor (5th Fl.) Session Panel: Who Builds the Frontend Now? (And What Breaks When They Do) Luca Mezzalira, James Hall, Danielle An, Ivan Zarea This panel is for frontend and mobile engineers trying to make sense of what's actually happening versus what's hype. Practitioners from platform infrastructure, observability, serverless architecture, digital consultancy, and mobile development will share what they're seeing on the ground and debate what it means. 14:45 Windsor (5th Fl.) Session GenAI Architecting AI Driven Game Creation: From Research to Production Scale Systems Danielle An Principal Engineer / GenAI Architect @Meta, Ph.D. with 15 years of professional experience in film, MR and gaming I spent the last two years as an architect of the next mobile gaming platform at meta, with the ambition that anyone can create, share, and play AI-generated mobile games. What I learned is that the hard part is no longer the generation — it's everything that happens after. 15:55 Mountbatten (6th Fl.) Session Running AI at the Edge: Running Real Workloads Directly in the Browser James Hall Founder and Director @Parallax, Author of jsPDF Running AI today often means choosing between Anthropic or OpenAI, and accepting the cost and privacy concerns that come with shipping data to third-parties. What if there was another way?