A Half-Century History of SQL
Today, most major relational databases support SELECT queries. That common feature across products can make it seem as though SQL was the standard database language from the start.
But half a century ago, the database field was full of proprietary interfaces that could not communicate with one another. SQL became an enduring industry standard not because it was the only option, but through decades of work: mathematical theory, rigorous systems engineering, commercial competition, and international standardization.
That history shows how one language broke through barriers between hardware platforms and vendors, reshaping the software industry’s approach to data.
First, Distinguish the Relational Model, SQL, and Database Products
To understand how SQL evolved, we first need to separate three concepts that are often conflated: the relational model, the SQL query language, and database products.
+-----------------------+
| Relational Model |
| (Data Concept) |
| Codd (1970) |
+-----------------------+
|
v
+-----------------------+
| SQL Language |
| (Query Syntax) |
| SEQUEL (1974) |
+-----------------------+
|
v
+-----------------------+
| Database Products |
| (Storage Engine) |
| System R, Oracle, Db2 |
+-----------------------+
In 1970, IBM researcher E. F. Codd published the landmark paper “A Relational Model of Data for Large Shared Data Banks”. At the time, hierarchical and network databases dominated, and applications had to know the physical location of data on disk and the pointer paths used to reach it.
Codd’s central idea was data independence: users should be concerned with the logical relationships among data, without depending on the underlying storage structure. A change to the storage format or an index should not require a change to the application’s query logic.
Grounded in relational algebra and relational calculus, the paper established a theoretical framework. It did not yet define the text-based query syntax familiar to engineers today.
In fact, the relational model and SQL are not quite the same. A mathematical relation is a set and contains no duplicate elements. By default, SQL queries retain duplicate rows; removing them requires DISTINCT. In other words, SQL draws on the relational model without reproducing it exactly.
SEQUEL Is Born: Querying Data with English Keywords
Once the relational model was in place, the next step was to design a practical language for working with it. In 1974, IBM researchers Donald Chamberlin and Raymond Boyce published “SEQUEL: A Structured English Query Language”.
The researchers wanted to avoid the difficult mathematical notation of formal logic, such as quantifiers and set operations. They instead used keyword patterns resembling structured English, so engineers and business analysts could express what they needed:
SELECT c.name, o.amount
FROM customers AS c
JOIN orders AS o ON o.customer_id = c.id
WHERE o.amount >= 1000;
The appeal of a declarative language is that users state what they want and leave how to the system.
The researchers also conducted a small experiment to see whether SEQUEL was easier to learn. In 1975, Phyllis Reisner, Boyce, and Chamberlin recruited university students for a human-factors study comparing SEQUEL with SQUARE, a language that used mathematical subscripts.
After 12 to 14 hours of instruction, students had acquired a working grasp of both languages, and those without programming experience performed better with SEQUEL. Features such as GROUP BY remained difficult to learn.
IBM then built a prototype SEQUEL interpreter on top of XRM, a relational memory system, to test how well the language could express queries. Because SEQUEL was already a registered trademark of a subsidiary of Hawker Siddeley, a British aerospace firm, the team dropped the vowels and adopted the name we know today: SQL.
System R: Testing Feasibility in a Real System
The prototype interpreter on XRM showed that the syntax could work, but the industry still faced a major technical question in the 1970s: if users only specified what they wanted, could a computer find an efficient access path in a reasonable amount of time? Even elegant syntax would remain in the laboratory if performance was poor.
To tackle that engineering problem, IBM launched an experimental project called System R. Its aim was to test the feasibility of a relational database with real disk storage and multiple users. The 1976 system paper covered views, permissions, integrity, transactions, logging, and recovery. It also stated clearly that System R was a research vehicle, not a planned product.
System R’s most consequential breakthrough was the cost-based query optimizer developed by Patricia Selinger and colleagues in 1979. Consider the earlier order query: finding the relatively few orders that meet the amount threshold before matching them to customers could cost vastly less than reading both entire tables before comparing them.
Using data statistics, the optimizer estimated CPU and I/O costs and chose the access paths and join order for the user. It did not guarantee the theoretically best plan every time, but it transferred the burden of deciding how to execute a query from the person writing it to the system.
That made data independence useful in practice. As the number of orders grew, an administrator could add an index and the database could adjust its execution plan without changing a line of application query logic. Of course, indexes, statistics, and each vendor’s optimizer still affect performance. SQL does not guarantee that a query, once written, will never need tuning.
The team also developed authorization and view mechanisms, including the ability to grant and revoke access to tables, so multiple users could safely share the same data.
Businesses needed SQL in programs that ran routinely, not just for one-off queries at a terminal. The System R team demonstrated that SQL could serve both uses: ad hoc terminal queries and transactions embedded in PL/I or COBOL programs that ran repeatedly. For the latter, the system could parse the SQL and select an access path before execution, avoiding that work each time the transaction ran.
System R’s success eased industry doubts about performance and laid the groundwork for commercial SQL products.
Commercialization: The Market Moved Ahead of the Standard
Entrepreneurs recognized a major business opportunity in System R’s research. IBM, meanwhile, was cautious about commercialization because of the profits from its existing hierarchical mainframe database, IMS, and published its research papers openly.
In 1979, Relational Software (later Oracle), co-founded by Larry Ellison, Bob Miner, and Ed Oates, used those public papers to launch Oracle V2, the world’s first commercial SQL database, ahead of IBM’s own product.
Facing market competition, IBM introduced SQL/DS in 1981 and launched its flagship mainframe product, Db2, in 1983.
The era followed a distinctive pattern: commercial adoption ran well ahead of formal standardization. Vendors built their own SQL engines from the research papers, greatly accelerating adoption while also sowing the seeds of syntax extensions and differences between implementations.
A Competitor’s Choice: Ingres Moves Away from QUEL
SQL was not the only contender in the early exploration of relational databases. Ingres, developed at the University of California, Berkeley around the same time, used a different query language called QUEL.
| Aspect | SQL (IBM / Oracle) | QUEL (Ingres / Berkeley) |
|---|---|---|
| Design approach | Structured keyword patterns resembling English | A rigorous, orthogonal tuple relational calculus (operations based on rows) |
| Representative systems | System R, Oracle V2, IBM Db2 | Ingres, POSTGRES (early versions) |
| Historical trajectory | Established as an international standard by ANSI/ISO | Gradually gave way from the 1980s onward; products eventually switched to SQL |
Many technical experts considered QUEL—built on the rigor of tuple relational calculus, a form of logic that operates on individual rows—simpler and more semantically consistent in design than early SQL. But as IBM and Oracle successfully promoted SQL commercially, buyers and the software ecosystem increasingly favored it.
Historical records reveal the practical side of this shift. An internal Ingres sales document from around 1985 shows that although the development team firmly believed QUEL was technically superior, Ingres had to offer a SQL interface in response to strong customer demand.
SQL co-creator Chamberlin described the same progression in an oral history: products that originally used QUEL first offered SQL as an alternative interface, then made it the primary one. He also emphasized that the user-facing language was converging, while each vendor still developed its own underlying technology.
By 1988, official Ingres documents explicitly described SQL as the industry standard. Similarly, POSTGRES, the predecessor of PostgreSQL, originally used PostQUEL before switching to SQL in 1995.
Ingres and POSTGRES added SQL interfaces in turn, showing that market demand and the software ecosystem had more influence on product choices than developers’ preferences about query-language design.
The Road to Standardization: From Common Usage to International Rules
With the market tilting further toward SQL, standards bodies began to establish formal specifications. Once standards were in place, SQL adoption expanded further:
- 1986: The American National Standards Institute (ANSI) published the first official SQL standard.
- 1987: The International Organization for Standardization published ISO 9075:1987, establishing SQL as an international standard.
- 1987: The US National Institute of Standards and Technology (NIST) approved FIPS 127, bringing SQL conformance into federal procurement testing and giving it substantial institutional backing.
FIPS 127 also brought the standard into conformance testing. A list of validated products published by NIST in 1995 included SQL products tested against the federal standard. Vendors could no longer simply say, “We support SQL.” They also had to answer, “Which specification do we meet, and how can we prove it?”
The standard continued to evolve. SQL-92 established much of today’s most commonly used core syntax and divided conformance into three levels: Entry, Intermediate, and Full. SQL:1999 moved to a conformance framework based on core features and optional packages. SQL:2023 continues to address modern needs, including JSON and property-graph queries.
Chamberlin noted in an interview that the standard created a shared market for tools, books, and courses. Publications, training, ORM tools, and developer skills could thus be used across systems, substantially reducing the cost of learning and integration.
Conclusion: A Shared SQL Standard, Different Implementations
Even with a SQL standard, databases do not all use precisely the same syntax. Retrieving just the five highest-value orders can involve two different ways to write the query:
-- A dialect form commonly used with PostgreSQL and MySQL
SELECT id, amount
FROM orders
ORDER BY amount DESC
LIMIT 5;
-- The form defined by the SQL:2008 standard (also supported by PostgreSQL)
SELECT id, amount
FROM orders
ORDER BY amount DESC
FETCH FIRST 5 ROWS ONLY;
According to the official PostgreSQL documentation, LIMIT is a product extension, while FETCH FIRST is standard SQL:2008 syntax. Vendors implemented their own versions before the standard emerged and have continued to add features. Existing programs depend on those forms, so dialects do not disappear when a standard is published. Differences extend to connections and error handling, too: query syntax alone cannot tell you whether an entire application can move to another database unchanged.
From Codd’s model and System R’s engineering validation to commercial products and international standards, SQL gave developers one basic language for communicating with different databases. That is its most significant achievement over half a century. When switching databases, developers still need to check whether their code relies on the standard core, optional features, or a particular product’s extensions. SQL unified the basic query language; it did not guarantee that applications could switch databases without changes.
NOTE
Further reading: OCI 101: The Common Language of the Container World