Bottom Linear Gradient  Lines image

Article

5 min

read

What Is “Prompt and Pray”?

And why it breaks voice AI agents

Dani Plicka Headshot

Dani Plicka

Content Marketing Manager

What is prompt and pray, and how to fix it

In this article

Share

Angular Gradient Image

Build it free.

Create a space and ship your first call flow in minutes.

Subscribe

Tags

AI Agents

Voice AI

PUC

Latency & Performance

“Prompt and pray” is a term for building an AI agent by writing detailed instructions into a prompt and hoping the model follows them. There is no code enforcing the instructions. There is no fallback if the model interprets a caller's request differently than intended. The prompt is the entire safety mechanism.

The term is used across the AI engineering community to describe why agents that look impressive in a demo start failing once real callers, real edge cases, and real pressure show up.

Why teams rely on “prompt and pray”

“Prompt and pray” is not a mistake teams make on purpose. It’s the natural starting point. A team writes a system prompt describing the agent's role, adds a few tools, tests it, and it works. Nothing in that process signals a problem.

The problem shows up later, once the agent meets a request nobody anticipated. A well-written prompt raises the odds the model behaves correctly. It does not guarantee it, because a prompt is guidance, not enforcement. The model is still free to interpret, improvise, or drift.

Why relying on prompts works at first and fails at scale

Relying on prompts tends to work for narrow, low-stakes tasks: reading back business hours, answering a simple FAQ, taking a message. The cost of a wrong answer is low, and the model has few ways to go wrong.

It breaks down when three things happen at once: the number of tools grows, the conversation gets longer, and the stakes of a wrong action rise. Tool selection accuracy degrades once an agent has more than a handful of tools to choose from at any one time. 

Long conversations dilute early instructions, and real phone calls are long, full of interruptions, corrections, and callers changing their minds. A model's attention to any one instruction competes with everything said since. And once an agent can issue a refund, transfer a call, or move money, a misinterpreted instruction stops being a minor annoyance and becomes a serious incident.

What “prompt and pray” gets wrong about the model

The main mistake you can make is treating the prompt as a contract when it is closer to a strong suggestion. A prompt describes what the model should do. It does not remove the model's ability to do something else if the conversation drifts far enough from the examples the prompt anticipated.

This flaw is just a property of how language models generate output: probabilistically, based on everything in context, not by executing a fixed set of rules. A system that depends entirely on prompt compliance for safety or correctness is exposed to that gap by design. Adversarial callers probe boundaries the prompt author never anticipated, and a model update can change behavior without warning – conditions a prompt has no way to hold up under.

The failure is often quiet rather than dramatic. For example, one team's pizza ordering agent could start telling callers that several menu items are sold out, though nothing is out of stock and no one had told it so. Nothing crashed. The model simply inferred a fact from nothing and stated it with confidence, which is a harder failure to catch than an outright error.

The fix is not a better prompt

The instinct when prompting fails is to write a longer, stricter prompt: more edge cases, more warnings, more capitalized "DO NOT" instructions. This only treats only the symptom because the prompt was never the layer that should have been carrying that weight.

The fix is to stop asking the prompt to do a job that software does more reliably. Tool availability, data access, and state changes can be handled by code that runs the same way every time, instead of by instructions the model has to remember and re-apply on every turn. 

The model keeps the job it’s good at: understanding intent and producing natural language. The software handles the job of guaranteeing the outcome.

The real replacement for over-prompting: Programmatically Governed Inference

This shift has a name: PGI (Programmatically Governed Inference), an architecture where the model requests actions and software decides what actually happens. We also refer to this as System-Directed AI.

PGI is the direct answer to “prompt and pray”. Instead of asking the prompt to carry every safety guarantee an agent needs, PGI removes that burden from the prompt entirely. The model can only see the tools relevant to the current step, can only navigate to steps it is explicitly allowed to reach, and can only request actions that a handler then verifies and executes. 

A prompt can be ignored, misread, or drift under pressure. But when code controls the agent, a tool that was never in the model's schema cannot be called by accident, and an action a handler never approves cannot happen at all. That’s the difference between hoping an agent behaves and building one that cannot behave otherwise.

Signs an agent is stuck because of “prompt and pray”

A few patterns are reliable indicators:

  • Instructions from early in a conversation stop being followed as the conversation gets longer.

  • The team's response to a new failure is always "add another line to the prompt."

  • The agent has a tool for something it should never be allowed to do, gated only by an instruction telling it not to use that tool.

  • Nobody can say with certainty what the agent will do in a scenario that has not been tested yet.

Any one of these is a sign the prompt is carrying enforcement weight it was never designed to hold.

Are you building the future of voice AI? Join our developer community on Discord to meet our team and others who are building it, too.

Top Linear Gradient  Lines image

Frequently asked questions

Frequently asked questions

The questions we hear most, answered.

The questions we hear most, answered.

What does "prompt and pray" mean?

It describes building an AI agent by relying entirely on prompt instructions for safety and correctness, with no code enforcing those instructions.

Is prompt and pray the same as prompt engineering?

No. Prompt engineering is the practice of writing effective prompts. Prompt and pray describes an architecture where prompt engineering is the only safety mechanism an agent has, with nothing enforcing the result.

Why does prompt and pray fail as agents scale?

Longer conversations dilute early instructions, larger tool sets reduce selection accuracy, and higher-stakes actions raise the cost of any single misinterpretation.

What replaces prompt and pray?

Architectures that separate language generation from action execution, often described as system-directed AI, with PGI (Programmatically Governed Inference) as one specific implementation of that approach.

Can a better prompt fix prompt and pray?

 A better prompt can reduce the frequency of mistakes. It cannot eliminate the underlying issue, since the prompt is still guidance rather than enforcement.

Bottom Linear Gradient  Lines image

The Communications Stack for What's Next

APIs built for speed. Infrastructure built for scale. AI built in from day one.

The Communications Stack for What's Next

APIs built for speed. Infrastructure built for scale. AI built in from day one.

The Communications Stack for What's Next

APIs built for speed. Infrastructure built for scale. AI built in from day one.

The Communications Stack for What's Next

APIs built for speed. Infrastructure built for scale. AI built in from day one.