The History of Ansible, Part 1: Agentless Design, YAML, and the Red Hat Acquisition
On the evening of February 23, 2012, Michael DeHaan made the first commit in a new Git repository. Its message was one word: “Genesis.” The codebase had just 12 modules at the time. Seven years later, the module directory in version 2.9 had grown to 3,908 files.
Starting as one person’s weekend project and backed by a single 100 million within three years. But two decisions made at the project’s outset defined its technical direction: abandoning resident daemons in favor of temporary SSH connections, and using YAML throughout instead of inventing a dedicated DSL.
These choices gave Ansible an exceptionally low barrier to entry and helped it break through when Puppet and Chef dominated the field. Over time, though, they also became a substantial source of architectural debt. DeHaan left the project in early 2015, and the company was acquired that October. In the years that followed, he acknowledged more than once that Ansible had strayed from his original vision of an extremely simple tool.
Before 2012: Cobbler, Func, and the Gap Between Them
Ansible did not emerge from nowhere. It was DeHaan’s third attempt to solve the same kind of systems operations problem.
In 2005, DeHaan joined Red Hat’s emerging technology group and built Cobbler, a provisioning tool that automated operating system installation on bare-metal machines through PXE network boot. Cobbler handled the process of bringing a server into existence. But DeHaan soon noticed an operational gap between the completion of OS installation and the point when a configuration management tool formally took over.
To fill that gap, DeHaan and Seth Vidal, among others, built the remote task execution framework Func. Func used a conventional architecture of resident daemons and PKI certificates: a central authority issued certificates so machines could authenticate one another. It even ported Puppet’s certificate management mechanism outright.
Building that system firsthand showed DeHaan just how fragile resident daemons and certificate trust chains could be in real data centers.
After leaving Red Hat, DeHaan briefly joined Puppet Labs. That experience became a catalyst for starting over on a tool of his own. He wanted a tool that did not require an agent on every machine or ongoing certificate maintenance, as Puppet did, yet could push commands from a controller directly to remote machines, as Capistrano, an SSH deployment tool popular in the Ruby community, did.
+----------------------+
| Bare-metal Provision |
| (Cobbler, 2005) |
+----------------------+
|
v
+----------------------+
| Remote Execution |
| (Func, Seth/DeHaan) |
+----------------------+
|
|--> [ The Missing Gap ]
| (Ansible born here)
v
+----------------------+
| Config Management |
| (Puppet / Chef) |
+----------------------+
The configuration management market was then led by Puppet (2005) and Chef (2009). Both broadly followed CFEngine’s theory of desired-state convergence: a central master stored configuration, while an agent installed on each node periodically pulled it and corrected drift.
Further reading: The Rise and Fall of IaC and Terraform: A Thirty-Year History
This architecture was compelling in theory, but in practice it came with three substantial upfront costs:
- Installing agents on nodes: Every target machine needed an agent and its dependencies installed in advance. In a restricted environment or on a brand-new machine, even logging in to install one could be difficult.
- Certificate maintenance: The master issued a certificate for each node, bound to its hostname and an expiration date. Inconsistent DNS resolution, unsynchronized clocks (NTP), or an old certificate left behind after reinstalling a machine could break authentication and disconnect a node, often when “nothing had changed.”
- Learning a dedicated language: Operators had to learn Puppet’s declarative DSL or become proficient in Ruby syntax for Chef.
DeHaan later recalled that resolving certificate and network configuration problems alone could take days. If an automation tool takes three days of troubleshooting before it can run, much of its potential benefit disappears.
“I wanted automation to look like a grocery list.”
Two Technical Bets: Agentless Operation and YAML
By late February 2012, a prototype of Ansible had taken shape in just two weeks. Soon after the first commit, DeHaan announced the project to early users on the Cobbler mailing list.
Ansible’s first major bet was to do away with resident agents entirely and use a pure controller-driven push model.
[ Controller (Laptop / CI) ]
|
| Temporary SSH / Python script execution
v
+---------------+
| Target Host | (No Daemon, No PKI, Python only)
+---------------+
No resident process is left behind on managed nodes. Ansible uses the SSH channel already available on the system (or WinRM for Windows) to push a temporary Python script to the remote host, run it, collect the JSON result, and clean up afterward.
That decision broke with the prevailing model. Ansible does not rely on a daemon to check periodically for drift and correct it automatically. A playbook, a list of tasks describing the desired configuration, enforces that state only when an operator or CI triggers a run. A single run still retains module-level idempotence: running the same configuration again does not change items that already match the desired state.
| Aspect | Master-agent model (Puppet / Chef) | Agentless push model (Ansible) |
|---|---|---|
| Node prerequisites | Install an agent and its dependencies on each target in advance | Only SSH and Python on the target |
| Trust and security | Maintain a separate master/node PKI certificate chain | Reuse existing SSH keys and sudo permissions |
| Learning curve | Learn a dedicated DSL or become proficient in Ruby | Write an intuitive YAML declaration |
| Execution and convergence | Background daemon polls regularly and corrects drift automatically | Controller pushes on demand; each run preserves idempotence |
| Time to first working run | Network and certificate issues may take hours or days to resolve | An engineer can run it from a personal laptop within 30 minutes |
Early contributor Seth Vidal, the author of yum, gave DeHaan a piece of advice that stuck:
“If people can’t get it working in about 30 minutes, they’ll leave.”
DeHaan made that a guiding principle — in his own words, “you have to let them succeed during a lunch break” — and spent roughly 40 percent of his early time writing clear, straightforward documentation.
The second major bet was to set aside a custom DSL and define playbooks entirely in YAML.
Many people see YAML as a deliberate choice for readability. Years later, DeHaan admitted a more practical reason: as a solo developer, he did not want to write and maintain a custom parser.
Using an established format meant existing tools could handle syntax parsing and editor support. System administrators did not need to learn Ruby or a special-purpose language; they could start by writing simple indented key-value pairs.
Both bets paid off quickly. RedMonk’s 2015 figures showed Ansible gaining about 200 forks per month, roughly twice Salt’s pace, with community activity pulling ahead through 2014 while Puppet and Chef largely held steady. Its module count rose from 57 in version 1.0 to 400 by the time DeHaan left, roughly a sevenfold increase.
Commercialization: Open Core, $6 Million in Funding, and the Red Hat Acquisition
As the project’s popularity surged, DeHaan, Saïd Ziouani, and others founded AnsibleWorks in March 2013. The company was later renamed Ansible, Inc.
Ansible followed an open-core business model. The automation engine was entirely open source under GPLv3, and its license never changed. Ansible Tower, a web management platform offering role-based access control (RBAC), audit trails, and scheduling, was commercial and closed source; it was only later released as the AWX project.
The CLI tool and modules engineers used day to day were free and open source. Larger enterprises that needed governance and compliance records were the paying customers.
In August 2013, the company raised a $6 million Series A round led by Menlo Ventures. It was the only funding round Ansible raised in its lifetime. Compared with competitors that spent heavily on marketing, Ansible grew organically because it was easy to adopt.
On October 16, 2015, Red Hat announced its acquisition of Ansible, Inc. Media reports put the price between 150 million. That made the acquisition price more than 16 times the total funding raised.
The Creator’s Departure: Module Sprawl, a Three-Year Noncompete, and JetPorch
Yet as the project neared its commercial peak, DeHaan chose to leave the company and the project in early 2015.
After his departure, he signed a three-year noncompete agreement that kept him out of systems automation from 2015 to 2018. Those were the very years when Docker containers, Kubernetes, and modern CI/CD were reshaping the infrastructure landscape.
In later interviews, DeHaan spoke openly of his disappointment with the project’s direction:
“There were probably around 400 modules when I left. There really should have been fewer. We could have held the line … but it grew in too many directions to stay focused.”
His sharper criticism addressed code quality and what the tool had become:
“It’s a hard one. I’ve said many times that the code in all my other projects is, overall, better than that one.”
“It was supposed to be ‘here’s the list of packages I want installed, here’s the list of files I want copied over.’ [What users ended up writing] wasn’t the direction I wanted.”
The growth in modules was no accident. From roughly 400 modules in 2015, the module directory grew to 3,908 files in version 2.9 as the commercial ecosystem expanded. Every cloud provider and hardware vendor wanted its own modules in the core repository, tying obscure modules to the same release cycle as core ones.
Early official attempts to split the repository did not succeed. Only years later did the Collections architecture allow modules to become separate packages with their own release schedules.
After the noncompete period ended, DeHaan eventually launched a new open-source project called JetPorch in July 2023. He wrote the underlying engine in Rust to improve performance, but still chose a dialect of YAML for the interface.
The landscape had changed, however. In late December 2023, DeHaan announced the end of JetPorch development. The 2012 pain point of configuring physical machines one by one over SSH had diminished substantially as cloud-native systems and containers took over much of that work. DeHaan returned with a more mature engineering design, only to find that the arena had shifted.
The Cost of the Bets: An SSH Performance Ceiling and YAML Stretched Too Far
Agentless operation and YAML lowered Ansible’s barrier to entry in its early days. As the number of managed machines and the complexity of automation grew, however, they imposed costs in performance and maintainability.
The first cost was a connection-performance ceiling caused by the absence of resident agents.
Ansible’s default execution model iterates over tasks and hosts (the per-task, per-host loop).
For each task, the controller packages the module code, transfers it to a temporary directory on the remote host, starts remote Python to run it, retrieves a JSON result, and removes the temporary files.
Ansible Default Task Loop (Per Host, Per Task):
+----------+-->+---------------+-->+---------+-->+-------+
| SSH Conn |-->| Transfer Code |-->| Execute |-->| Clean |
+----------+-->+---------------+-->+---------+-->+-------+
^ |
+------ Next Task Loop -------------------------+
According to a quantitative analysis by the performance optimization tool Mitogen, starting remote Python processes and importing modules alone adds 300 to 800 milliseconds per playbook step, before the cost of establishing SSH connections is even counted. For a 100-task playbook, those overheads alone amount to 30 to 80 seconds.
To remain compatible with the default sudo configuration on some Linux distributions, Ansible also disables pipelining by default. Pipelining speeds execution by sending code directly over the connection instead of writing a temporary file. The default number of concurrently managed hosts, forks, is conservatively set to 5. At a scale of hundreds of machines, these choices create a noticeable performance ceiling.
The project briefly tried spinning up temporary background processes on remote hosts to accelerate execution, through modes such as fireball and accelerate. Those modes were all retired because they conflicted with the agentless promise of leaving no daemon on managed hosts.
This helped give rise to the third-party Mitogen plugin. By keeping a Python interpreter running in memory on target machines, Mitogen claims speedups of 1.25 to 7 times. But the need for a third-party plugin to change a core tool’s internal execution model to make it fast illustrates the limits built into the original architecture.
The second cost was forcing a data serialization language to act like a programming language.
YAML was designed to express relatively simple data structures. As operational logic grew more complex, conditions (when), loops (loop), exception handling (block/rescue), and variable assignment (register) were pressed into YAML fields, while expressions relied on Jinja2 template strings:
- name: Complex logic pieced together across YAML and Jinja2
ansible.builtin.shell: /usr/local/bin/deploy.sh
register: deploy_result
loop: "{{ target_services | selectattr('enabled') | map(attribute='name') | list }}"
when:
- inventory_hostname in groups['production']
- deploy_result.rc is not defined or deploy_result.rc == 0
failed_when: "'CRITICAL' in deploy_result.stderr"
A single playbook now mixes two syntaxes, with errors coming from two separate systems: bad indentation causes a YAML parsing failure; faulty logic produces a Jinja2 undefined-variable error. There is also no type system or native debugger, and IDEs struggle to perform precise static analysis.
DeHaan himself put it this way:
“It looks like a bad programming language because it was never supposed to be a programming language.”
Even so, pyinfra, a comparable tool that lets users write configuration in plain Python and expresses logic more rigorously, has only a fraction of Ansible’s GitHub stars. In terms of adoption, a low barrier to entry appears to matter more than the expressive power of a programming language.
Conclusion: The Victory and Cost of a Low Barrier to Entry
Ansible’s first act showed how a tool with a low cognitive burden could prevail over a more complex architecture. An agent-free design and approachable YAML directly answered early-2010s operators’ frustration with PKI certificates and dedicated DSLs, carrying Ansible, on just one funding round, all the way to an acquisition.
But turning configuration declarations into program logic and using temporary SSH connections for large-scale workloads also planted the seeds of technical debt that would constrain its growth.
Part 2 will explore how Ansible repositioned itself in the modern cloud-native ecosystem as Terraform moved infrastructure toward immutability, Kubernetes replaced long-lived servers for many workloads, and Ansible itself underwent the major split into Collections.