Preprint / Version 0

Surfer 2: The Next Generation of Cross-Platform Computer Use Agents

Authors

  • Mathieu Andreux
  • Märt Bakler
  • Yanael Barbier
  • Hamza Benchekroun
  • Emilien Biré
  • Antoine Bonnet
  • Riaz Bordie
  • Nathan Bout
  • Matthias Brunel
  • Aleix Cambray
  • Pierre-Louis Cedoz
  • Antoine Chassang
  • Gautier Cloix
  • Ethan Connelly
  • Alexandra Constantinou
  • Ramzi De Coster
  • Hubert de la Jonquiere
  • Aurélien Delfosse
  • Maxime Delpit
  • Alexis Deprez
  • Augustin Derupti
  • Mathieu Diaz
  • Shannon D'Souza
  • Julie Dujardin
  • Abai Edmund
  • Michael Eickenberg
  • Armand Fatalot
  • Wissem Felissi
  • Isaac Herring
  • Xavier Koegler
  • Erwan Le Jumeau de Kergaradec
  • Aurélien Lac
  • Maxime Langevin
  • Corentin Lauverjat
  • Antonio Loison
  • Avshalom Manevich
  • Axel Moyal
  • Axel Nguyen Kerbel
  • Marinela Parovic
  • Julien Revelle
  • Guillaume Richard
  • Mats Richter
  • Ronan Riochet
  • María Santos
  • Romain Savidan
  • Laurent Sifre
  • Maxime Theillard
  • Marc Thibault
  • Ivan Valentini
  • Tony Wu
  • Laura Yie
  • Kai Yuan
  • Jevgenij Zubovskij

Abstract

Building agents that generalize across web, desktop, and mobile environments remains an open challenge, as prior systems rely on environment-specific interfaces that limit cross-platform deployment. We introduce Surfer 2, a unified architecture operating purely from visual observations that achieves state-of-the-art performance across all three environments. Surfer 2 integrates hierarchical context management, decoupled planning and execution, and self-verification with adaptive recovery, enabling reliable operation over long task horizons. Our system achieves 97.1% accuracy on WebVoyager, 69.6% on WebArena, 60.1% on OSWorld, and 87.1% on AndroidWorld, outperforming all prior systems without task-specific fine-tuning. With multiple attempts, Surfer 2 exceeds human performance on all benchmarks. These results demonstrate that systematic orchestration amplifies foundation model capabilities and enables general-purpose computer control through visual interaction alone, while calling for a next-generation vision language model to achieve Pareto-optimal cost-efficiency.

References

Downloads

Posted

2025-10-24