Chatbots did not begin with ChatGPT, and their history is more than a sequence of famous demos. Each generation changed who wrote the responses, how much context the system could use, which channels it could handle, and what teams had to operate. Following those changes explains how scripted text programs became today's tool-using, multimodal conversational systems, and why a production voice agent is still much more than a language model.
The history of chatbots in one view
The history of chatbots is a sequence of shifting engineering boundaries. Early systems stored conversation in rules. Statistical systems learned narrow language patterns from labeled data. Neural models learned to generate replies. Transformers made broad pretraining practical. Large language models then combined that pretraining with instructions, examples, retrieval, and tools. Current conversational systems add live audio, images, and application state.
Each advance expanded what a person could say. It also introduced a new class of failure and a larger production surface.
| Era | Interaction model | Training and control | Context | Main channels | Operating burden | Typical failure |
|---|---|---|---|---|---|---|
| Rule-based chatbots, 1960s–1990s | Match a pattern and return a scripted response | People write rules, templates, and priorities | Current input plus a few stored variables | Text terminals and desktop programs | Author and maintain every path | Unseen wording falls through or triggers the wrong rule |
| Statistical natural-language processing (NLP), 1980s–2000s | Classify intent, extract fields, retrieve an answer | Train probabilistic models on labeled speech or text; keep dialogue policy explicit | Slots and session state | Text, telephone menus, early speech interfaces | Collect labeled data and integrate back-end services | Recognition, intent, and slot errors compound |
| Neural dialogue, 2014–2017 | Generate the next response from preceding text | Train encoder-decoder networks end to end on conversation pairs | A short history compressed into a hidden state | Mainly text | Curate large datasets and serve models | Generic, inconsistent, or repetitive replies |
| Transformers and pretraining, 2017–2020 | Attend across tokens and predict the next token | Pretrain broadly, then fine-tune for a task | A larger token window | Mainly text | Supply compute, data, fine-tuning, and evaluation | Fluent output can still be false or incoherent |
| Large language model (LLM) assistants, 2020–2023 | Follow open-ended instructions in dialogue | Pretraining plus instruction tuning and human feedback | Prompt, chat history, and in-context examples | Web and app chat, then APIs | Manage prompts, safety, cost, versions, and evaluation | Hallucination, bias, inconsistent constraint following |
| Grounded and tool-using agents, 2020s | Retrieve information and propose actions in external systems | Add retrieval, tool schemas, policies, and task-specific evaluation | Model window plus documents, tool results, and application state | Chat, apps, messaging, and APIs | Secure tools, handle retries, preserve state, and trace outcomes | Wrong retrieval, wrong arguments, unsafe or duplicate actions |
| Multimodal real-time systems, 2023 onward | Process and generate some mix of text, audio, images, and video | Multimodal pretraining, alignment, orchestration, and runtime policy | Conversation history plus live media and external state | Chat, web, mobile, phone, and connected devices | Operate streaming media, turn-taking, latency, tools, telephony, and monitoring | Missed interruptions, audio artifacts, hidden state errors, and hard-to-reproduce failures |
Chatbots, conversational AI, and production voice agents
These terms describe different scopes.
- A chatbot is the user-facing program that conducts a written or spoken exchange. It can be a fixed decision tree, a retrieval system, an LLM interface, or a mix of all three.
- Conversational AI is the wider technical category. It can include speech recognition, natural-language understanding, dialogue management, generation, retrieval, tools, speech synthesis, channels, and analytics.
- A production voice agent is a running system that takes part in live spoken conversations and completes defined work. It must handle streaming audio, turn-taking, interruptions, calls, tools, state, failures, and operational evidence within a tight time budget.
A model supplies language or audio intelligence. A runtime coordinates that intelligence with the rest of the product. Our AI agent runtime guide explains why session lifecycle, tool execution, policy, recovery, and telemetry remain separate production concerns.
1950: Turing frames conversation as a test
Alan Turing did not build a chatbot in 1950. He proposed the imitation game as a practical replacement for the vague question of whether machines can think. The setup used written questions and answers so the evaluator could judge behavior without seeing the participants. Turing's 1950 paper made conversation a visible test of machine intelligence.
That idea shaped public expectations for decades. It did not specify how a conversational program should learn, remember, or act. The first working systems answered those questions with rules.
1966–1970s: ELIZA and PARRY create the rule-based pattern
Joseph Weizenbaum's ELIZA accepted typed input, looked for keywords and patterns, transformed parts of the sentence, and selected a response from a script. Its best-known DOCTOR script used the open-ended style of a psychotherapist, which let the program return many statements as questions. The original ELIZA paper describes decomposition rules, reassembly rules, keyword priorities, and a limited form of memory.
ELIZA did not learn from conversations. Its control was legible because a person could inspect the script. That same design imposed a hard ceiling. A wording the rules did not cover produced a generic response, and apparent understanding could disappear on the next turn.
PARRY, built by psychiatrist Kenneth Colby in the early 1970s, added a more explicit model of internal variables and beliefs to simulate a person with paranoid thinking. Colby, Sylvia Weber, and Franklin Hilf documented the model in their 1971 Artificial Paranoia paper. This was a step toward stateful dialogue: the response depended on more than a single text pattern. The domain remained deliberately narrow, and the behavior still came from authored logic.
What changed: conversation became a software interface. Developers owned nearly all language and control by hand. Context meant keywords, recent phrases, and a small set of variables. The channel was typed text. The main operating cost was rule coverage, and the main failure was brittleness outside that coverage.
1980s–2000s: statistical NLP and messaging broaden the interface
Statistical methods changed how systems recognized language. Hidden Markov models treated speech as a sequence of unobserved states that produced observed acoustic signals. They became a foundation of speech recognition, as summarized in Rabiner's influential hidden Markov model tutorial. Statistical classifiers and language models also made it possible to rank likely words, intents, or responses from data rather than encode every variation manually.
This did not make rule-based chatbots disappear. The two approaches coexisted. A.L.I.C.E., developed from 1995, organized large collections of pattern-response rules in Artificial Intelligence Markup Language (AIML). Creator Richard Wallace's technical account of A.L.I.C.E. describes the program and AIML's development from that year. SmarterChild launched on AOL Instant Messenger in 2001 and later reached Yahoo! Messenger, where users could chat or request structured information such as weather and stock quotes, according to the Computer History Museum's account.
Distribution and integration were as important as model quality. A bot could now meet people in a familiar messaging channel and retrieve changing information. Dialogue systems also began to use explicit slots such as destination, date, or account number, then move through a designed task flow.
What changed: parts of language handling moved from authored patterns to probabilities learned from examples. Context became a session record with intents, slots, and back-end results. Text messaging and telephone speech widened access. Teams now had to label data and connect services as well as write dialogue. A speech-recognition error could corrupt intent classification, slot filling, and every step after it.
2007–2016: assistants combine speech, intent, and services
Voice assistants brought several existing technologies into one consumer experience. Siri grew from SRI's Cognitive Assistant that Learns and Organizes (CALO) research program. SRI spun out Siri, Inc. in 2007, Apple acquired the company in 2010, and Siri was integrated into the iPhone 4S in 2011. SRI's Siri history shows that the milestone was system integration: speech recognition, intent handling, service delegation, and a mobile interface working together.
Similar assistants and smart speakers made spoken requests routine. Messaging platforms also opened bot APIs to outside developers. Telegram launched its Bot API and platform in June 2015, and Facebook launched the Messenger Platform with bots and a Send/Receive API in April 2016. These channels brought decision-tree customer-service bots into brand conversations.
These systems worked well when an utterance mapped to a known intent and service. They struggled with requests outside their domain, follow-up references the state model had not anticipated, and conversational repair after a wrong recognition. Control stayed relatively strong because developers defined intents, fields, and action paths. Coverage was expensive because every new task needed examples, integrations, and dialogue logic.
2014–2017: neural networks learn to generate dialogue
Sequence-to-sequence models replaced separate hand-built language steps with an encoder that represented an input sequence and a decoder that generated an output sequence. The original sequence-to-sequence research focused on translation, yet the architecture soon reached dialogue.
In 2015, researchers trained a neural conversational model to predict the next sentence from previous sentences. It could learn from conversation data end to end with far fewer handcrafted rules. Its authors also reported a defining failure: lack of consistency.
The interaction model had changed from choosing a prepared reply to generating a reply token by token. Training data now controlled much of the behavior implicitly. Context was usually a short conversation compressed into a fixed representation or recurrent hidden state. The system could respond to wording no developer had anticipated, while debugging why it chose a particular reply became harder. Generic answers, repetition, contradictions, and drift were common.
2017–2020: transformers make broad pretraining practical
The 2017 Transformer paper replaced recurrent sequence processing with attention. Tokens could relate directly to other tokens in the input, and training could run with much more parallelism. The first results were in translation. The architecture later became the base for large pretrained language models.
Pretraining changed the development model. A team no longer needed to train every language capability from scratch for one bot. A general model could learn patterns from a broad text corpus, then be fine-tuned or prompted for a narrower job. Larger context windows let the model use more conversation history than earlier neural dialogue systems, although the history still had to fit inside a finite token budget.
The burden moved toward collecting and filtering training data, supplying compute, adapting the model, and measuring behavior across many possible prompts. Fluency improved faster than reliability. A model optimized to predict likely text had no built-in guarantee that a statement was true, current, authorized, or consistent with a business rule.
2020–2023: large language models and ChatGPT change the default interface
Scaling transformer language models expanded few-shot and in-context learning. The 2020 GPT-3 paper showed that one autoregressive model could attempt many language tasks from instructions and examples without a separate gradient update for each task.
Raw next-token prediction was still a poor control surface for an assistant. Instruction tuning and reinforcement learning from human feedback made responses more likely to follow a request and match human preferences. A 2022 instruction-tuning study also made the limitation clear: larger models were not inherently better at following user intent, and aligned models still made mistakes.
ChatGPT's public release in November 2022 packaged those advances in a simple, stateful chat interface. OpenAI's November 2022 announcement presented ChatGPT as a conversational model whose dialogue format supported follow-up questions and corrections. Its significance came from accessibility, open-ended usefulness, and distribution. It was one milestone in a line that already included decades of chatbots, statistical language systems, neural dialogue, and transformers.
Control now came from several layers: pretraining data, post-training, system instructions, conversation history, sampling settings, and application policy. The model could maintain a longer local thread and switch tasks without an intent schema. It could also invent facts, violate instructions, reproduce bias, and behave differently after a small prompt or model change. Operating a serious assistant therefore required evaluation, safety controls, version management, and monitoring.
2020s: retrieval and tools turn replies into actions
An LLM's parameters are an unreliable place to store changing business facts. Retrieval-augmented generation adds an external index and supplies selected documents at response time. The original RAG research framed this as a combination of parametric memory in the model and non-parametric memory in a retrievable corpus.
Tool use extends the loop again. The model can propose a structured call, receive the result, and continue. The ReAct study showed how interleaving reasoning and external actions could ground answers and support multi-step tasks.
This changed what context means. It now includes documents, customer records, tool results, workflow state, and permissions alongside recent messages. It also changed the dominant risk. A wrong answer is harmful; a wrong refund, booking, or record update creates a real side effect. Production teams must validate arguments, authorize each action, make retries safe, preserve authoritative state outside the model, and reconstruct failures across model and tool events.
2023 onward: multimodal and real-time systems converge
Current conversational systems can use several modalities within one session. Some combine text and image understanding. Others accept and generate audio tokens directly. AudioPaLM demonstrated a unified model for speech understanding and generation that combined linguistic knowledge with information carried by speech, including intonation. The AudioPaLM research marked a move beyond treating audio only as text waiting to be transcribed.
There are still two main ways to build a live voice conversation:
- A cascaded system connects voice activity detection, speech recognition, a text model or dialogue manager, and text-to-speech. Each component can be chosen, inspected, and controlled separately, but every boundary adds delay and another failure point.
- A speech-to-speech system models incoming and outgoing audio more directly. The Moshi paper describes parallel audio streams that support overlapping speech and interruptions while retaining time-aligned text. This can preserve vocal information and conversational timing, while making policy checks, transcripts, provider substitution, and component-level debugging harder.
Multimodality expands context from words to live signals. A system may need to decide whether a pause ends a turn, connect a spoken reference to an image, interrupt playback cleanly, and keep a tool result synchronized with the conversation. Channel behavior is part of correctness. A reply that works in text may be too slow, too long, or impossible to repair naturally on a phone call.
Why production voice agents are a separate engineering stage
A multimodal model can understand and generate audio. A production voice agent must keep a real conversation and a business process in a valid state.
That job includes media transport, telephony, turn detection, interruption handling, model and voice orchestration, tool execution, transfers, timeouts, concurrency, versioning, monitoring, and replayable evidence. The voice AI stack is a system of these layers. Changing one layer can alter latency, recognition, turn-taking, cost, and task success.
This is where the operating burden in chatbot history has landed. Early developers wrote nearly every response. Current teams spend more time defining boundaries and proving behavior: which context enters the model, which actions it may take, what must remain deterministic, how failures recover, and what evidence is available after a bad call.
At Dasha, we focus on this production layer. We help technical teams build and run voice AI agents through a managed runtime, REST APIs, and a web application for phone and web channels, with integrations, testing, monitoring, and call execution. Teams retain responsibility for business logic, connected systems, compliance, and production acceptance.
If your next step is choosing an operating model rather than studying its history, compare conversational AI platforms for production teams.



