Vivold Consulting
Product & Tech Updates

DeepMind Releases Gemini 2.5 Computer Use Model

Key Insights

DeepMind has unveiled Gemini 2.5, a model enabling AI agents to interact directly with graphical interfaces, performing tasks like form filling and navigation behind logins.

Stay Updated

Get the latest insights delivered to your inbox

Bridging AI and User Interfaces

  • DeepMind's Gemini 2.5 introduces a significant advancement by allowing AI agents to seamlessly interact with graphical user interfaces (GUIs).
  • This capability enables tasks such as form completion, scrolling, and operating within authenticated environments, broadening the scope of AI applications.

Implications for Automation and Accessibility


  • Enhanced Automation: Businesses can leverage Gemini 2.5 to automate complex workflows that involve GUI interactions, potentially reducing manual effort and increasing efficiency.

  • Improved Accessibility: The model's ability to navigate and operate within various interfaces could lead to the development of more accessible digital environments for users with disabilities.

Strategic Considerations


  • Competitive Edge: Organizations adopting Gemini 2.5 may gain a competitive advantage by streamlining operations and offering more responsive user experiences.

  • Integration Challenges: Implementing such advanced AI models requires careful integration with existing systems and consideration of ethical implications, particularly concerning user data privacy.
As AI continues to evolve, models like Gemini 2.5 highlight the importance of bridging the gap between artificial intelligence and human-computer interaction.

More in Product & Tech Updates

All Product & Tech Updates stories

Claude gets a 'Reflect' dashboard: Spotify Wrapped for your AI habit - and a masterclass in retention design

Anthropic launched Reflect, a beta dashboard (Free, Pro, and Max users with Memory on) that visualises Claude usage over 1-12 months - top topics, task types, peak hours - and coaches you via its 4D AI Fluency Framework (delegation, description, discernment, diligence), suggesting features like Projects or custom skills based on your patterns. It ships wellbeing controls (quiet hours, break nudges, reflection prompts) built with MIT Media Lab and Boston Children's Hospital experts, and excludes health-linked conversations entirely. TechCrunch's sharp read: beneath the mindfulness framing, Reflect showcases how much of your work runs through Claude - a retention play as much as a wellness one.

Gemini Spark lands on the Mac: Google's 24/7 agent starts working your local files

Gemini Spark, Google's 24/7 agentic assistant, is now available on Mac (beta, US-only, Google AI Ultra subscribers), where it can work directly with files on the computer - sorting and organising them, or turning a folder of invoices into a budgeting worksheet in Google Workspace. The update adds long-requested Google Tasks and Keep integrations plus third-party hooks into Canva, Dropbox, Instacart, OpenTable, and Zillow Rentals, real-time tracking of topics like stocks and breaking news, and - notably for builders - custom MCP support for wiring in your own apps. It puts Spark in direct competition with Claude Desktop, Microsoft Copilot, and OpenClaw for the desktop, where the real productivity (and governance) questions live.

L'Oreal brings Maybelline virtual try-on to ChatGPT

L'Oreal has announced a wide-ranging collaboration with OpenAI, unveiled at VivaTech 2026, that brings Maybelline's virtual makeup try-on directly into ChatGPT via L'Oreal's ModiFace AR technology. The deal spans consumer shopping tools, product discovery for brands like Lancome and Kerastase, advertising pilots (SkinCeuticals, CeraVe, Garnier), and R&D - including using OpenAI's GPT-Rosalind life-sciences model for skin-microbiome research. It lands as OpenAI reports ChatGPT at more than 900 million weekly users.