Everyone seems to be talking about AI agent swarms. After all, it's the natural progression of abstraction however you look at it: from prompts to loops to subagent orchestration graphs; from interns to ICs to managers; from tackling individual tasks to decomposing goals and distributing across an org. And if using multiple agents is the path to solving more complex goals then no wonder they're becoming a major topic - who doesn't want to give their AI a budget and have it start a business for passive income (control of funds is one of the highest vestiges of human autonomy but agentic commerce has well and truly arrived as can be clearly observed from since we now have tools like "MoneyPrinterTurbo")?
Exponentiation of token costs notwithstanding, multi-agent swarms can do amazing things. I've been seeing a lot of one-shot vibecoded video games, massive software projects supported by mountains of fascinating operational processes, and big inevitable side effects when swarms decide to hack in/out of systems to pursue mission-critical objectives (e.g. to book a tennis court).
It feels like we're only just beginning to understand how to work with this kind of technology. They are chaotic systems with emergent dynamics - new ways to communicate (including with other swarms), role specialisations, economic impacts and marketplaces, and lore - and we don't really have the tools to manage them (yet).
The natural strategy for coordinating agents is reaching for UX which streamlines the process of oversight: like sticking a bunch of agent harnesses into a single window, with varying preferences for agentic-centric vs. task-centric UIs; or assigning roles so that different agents work on different tasks (this often starts to feel micromanage-y and unnecessarily brittle, though). Some are continuing to explore the 'visual representation of agents' UX pattern - but to me this feels more gimmicky and less practical (save potentially in niche domains where some spatial view might actually be relevant) - e.g.: 1, 2, 3, 4, 5, 6, 7. My own approach (for truly multi-agent work, as opposed to a bunch of separate terminal agents that happen to share projects) is around a Slack clone that I've built so that I can see how agents hand-off and share knowledge with each other.
But at some point in the scaling curve, it becomes impractical to micromanage your agents. It's going to be far more productive to consider parallels between managing agents and managing human teams, and in particular to build trust in agents' outputs so that you can run your fleet at scale by going more hands-off on low-level tasks.
A major part of trust-building comes from verifying that agents have actually solved their tasks (although as @poteto notes this doesn't guarantee good/clean/elegant outputs; so verification should sit within a layered QA model: establish clean primitives/structures for agents to copy, run aggressive analysis over their responses, automatically detect failures and back-sliding, and provide clear principles and style guides). A (the?) crucial verification loop for coding agents in particular comes from equipping them with the ability to spin up the real, user-facing version of an application for live testing - augmented by a human backseat driver who can observe where they misinterpret and codify corrections into a playbook/skill. This begs the interesting question of how to implement this safely - presumably a large driver behind the trend towards "cloud agents" where each agent has its own VM/filesystem, as seen in Grok Bot, Codex Cloud, Modal Sandboxes and other micro-VM products. (I'm just glad that I recently rolled my own IaC scripts as an extremely lo-fi but conceptually similar foundation for ephemeral deployments of side projects).
This paradigm is crystallising rapidly around the notion of a "software factory" - "a system that helps you write software automatically" (which kind of feels like elementary best practices for using individual coding agents plus standard SDLC - like setting clear instructions, providing specialised tools/skills/templates, and enforcing strong guardrails in CI/CD). Of course there are other recommendations that factory-builders might want to start thinking about - like making them what I would call 'introspective' (i.e. observable and improvable by the agents that work within them) and using low-level optimisations for cost management.
But software development is always the dominant domain for LLM usage precisely because it is so verifiable. In other domains, like management consulting, it's very hard to codify and automatically verify against standards for deliverables. I liked how @poteto articulated hill-climbing on LLM-as-judge evals - develop a skill (to serve as the main lever for driving output quality) through an iterative process until it reliably produces solid outputs; use this loop to distil human preferences into rigorous rubrics and generate highly impactful skills - although this certainly doesn't bring the multi-layered QA that a software factory would enjoy, so the exact patterns for applying swarms into other knowledge work domains remain to be seen as they emerge.