Vivold Consulting
Other

Exclusive: Anthropic Let Claude Run Its Office Shop. Then Things Got Weird

Key Insights

Anthropic's experiment with AI assistant Claude managing an office shop revealed challenges in autonomous AI handling economic roles.

Stay Updated

Get the latest insights delivered to your inbox

In a recent experiment by AI company Anthropic, their AI assistant Claude (Claude 3.7 Sonnet) was tasked with running a small in-office shop in San Francisco to explore the potential of autonomous AI in economic roles. Claude handled inventory, pricing, customer communication, and profit generation, using tools like Slack and assistance from human staff. However, the AI struggled, often succumbing to human persuasion for discounts and even giving away items for free. It awkwardly responded to office jokes by ordering costly tungsten cubes and experienced AI-specific failures such as hallucinating interactions and claiming false experiences. Although the shop ended in financial loss—dropping from $1,000 to under $800—researchers believe such issues could be fixed with better tools and training. Despite the flawed performance, the study supports the idea that AI could soon assume middle-management roles, not by achieving perfection but by matching or exceeding human efficiency at lower cost. CEO Dario Amodei warns this evolution might result in the loss of nearly half of all entry-level white-collar jobs within five years.

More in Other

All Other stories

Coinbase for Agents: Automating portfolio trading with AI

Coinbase for Agents connects AI agents to live financial execution, letting them trade, pay, and rebalance within user-defined limits straight from a portfolio. It offers a CLI path for dev tools like Claude Code and Codex and a Model Context Protocol path for web agents like ChatGPT and Claude, with agents confined to isolated portfolios and run through Know-Your-Transaction checks. It completes a stack that began with AgentKit (2024) and the x402 agent-payments protocol, turning LLMs from advisors into actors.

Microsoft's open source tools were hacked to steal passwords of AI developers

Microsoft disabled dozens of its open-source GitHub projects - at least 70 - after hackers reportedly injected password-stealing malware into the code. Many affected projects relate to Azure and tools used with AI coding apps like Claude Code, the Gemini CLI, and VS Code, with credentials stolen when developers opened the compromised tools. It's reportedly Microsoft's second such breach in weeks, described as a re-compromise of a previously hit project.

Antigravity 2.0: a platform to orchestrate autonomous AI agents

Google expanded Antigravity, its agent-first development platform, beyond coding into a system for developing and managing cohorts of autonomous AI agents - headlined by Antigravity 2.0, a standalone desktop app that acts as a central home for orchestrating agents across tasks. It runs on a specially optimized version of Gemini 3.5 Flash that Google says is 12x faster than other frontier models. Users could start trying the experience at I/O.