Workflow AutomationOperationsZapierMaken8n

Automation Works in Test but Fails Live: 6 Causes

An automation that passes its test and fails live is almost never flaky. A test run differs from a live run in six specific inputs: the record, the trigger path, the volume, the timing, the saved state, and the error handling. Zapier, Make, and n8n each hide a different subset of these during testing, so the fix is to replay real records from run history rather than the editor sample.

Alexey YushkinFounder, GENERAL INFORMATICS3 min read

An automation that passes its test and fails on the first live run is almost never flaky. A test run differs from a live run in six specific inputs: the record it processes, the trigger path it enters through, the number of items, the timing, the saved state it carries, and what happens when it errors. Zapier, Make, and n8n each hide a different subset of these while you are building, so the fix is to find which of the six changed, then replay real records from run history instead of trusting the editor sample.

The forum answers to this question are usually one cause at a time: the Zap was off, the field was blank, the workflow was not activated. Each is right for one person. Below is the full list, so you can check all six in one pass instead of finding them in the order your customers do.

A test proves one record went through one path, once

A passing test is a narrow statement. It says that a specific record, chosen by you or by the editor, made it through every step in a single run that you started by hand and watched. Everything else about production is untested.

That is not a criticism of the tools. Test mode exists to let you map fields, and it does that well. The trouble starts when a green check mark in the editor gets read as "this works," when it means "this worked for that record, at that moment, on that path."

Here is how the six inputs differ between the test and the live run on each platform, as documented by the vendors in October 2026.

What changesIn the testLive on ZapierLive on MakeLive on n8n
The recordA sample you picked, usually recent and completeReal records with blanks and stray formattingSameSame
The trigger pathPulled on demand, or sent to a test URLPolling compares IDs against every ID seen since the Zap was turned onWebhooks run instantly, or wait in a queue if scheduled or offProduction URL, registered only when the workflow is published
The volumeOne itemMore than 100 new records in one poll are held by defaultMore than 300 webhook requests per 10 seconds return 429Production executions count toward the plan quota
The timingMinutes or hours after the record was createdPolls every 15 minutes on Free, 2 on Professional, 1 on TeamInstant, or on your scheduleInstant for webhooks
The saved statePinned data, cached samplesTesting an action step already created a real recordQueued items from while the scenario was offPinned data is ignored
Error handlingYou are watching the editorErrors go to Zap HistoryErrors follow your error handlersThe Error Trigger only runs for automatic executions

Read across a row, not down a column. The question is never "is Make worse than n8n," it is "which row bit me."

The record: the sample is the cleanest data you will ever see

Most first-run failures are in the first row. The test record is usually one you created while building, so you filled in every field and typed it carefully. Real records arrive from a web form on a phone, a CSV a colleague exported, or a CRM where half the fields are optional.

What breaks on the real record is predictable:

  • A field that was filled in the sample is blank live, and the step that uses it either errors or writes nothing. Blank can mean an empty string, a null, or the key missing from the payload entirely, which is why one fallback rarely covers all three. We wrote about the three kinds of empty and which fields should block a send.
  • An email or phone number has a trailing space, so the receiving app rejects it as invalid.
  • A name with an accent, an apostrophe, or a line break hits a step that only ever saw plain ASCII.

The diagnostic is mechanical. Open the failed run, open the test record, and put the inputs side by side. On Zapier that is Zap History; on n8n it is the Executions tab; on Make it is the scenario history. The field that differs is the cause in most cases, and you do not need to guess.

The trigger path: the live run did not enter the way the test did

The second row produces the most confusing failure, because nothing errors. The automation simply never runs.

Zapier polling triggers deduplicate by ID. Zapier's developer documentation says that when a Zap is first turned on, it retrieves existing data and stores each item's ID, and active Zaps then trigger only on IDs they have not seen. So the record you tested with, which existed before you turned the Zap on, will never trigger it live. Editing a record does not help either, because the ID is unchanged. A "new record" trigger fires for new IDs only, and an "updated record" trigger only works if it builds its key from the ID plus the update timestamp. The same documentation adds that the stored list is cleared when the Zap is turned off. Turning it back on repeats the initial call, so every record created while the Zap was off is filed as already seen and never triggers it.

n8n webhooks have two URLs. n8n registers the test URL when you select Listen for Test Event or Execute workflow, and registers the production URL only when you publish the workflow. If the sending app still points at the test URL, live events reach nothing once you close the editor. If it points at the production URL but the workflow is unpublished, the same. And because n8n does not display production data on the canvas, a run that did happen is easy to miss: it is in the Executions tab, not the editor.

Make webhooks queue when they cannot run. When a scenario is scheduled or switched off, Make stores incoming webhook data in a queue, capped at 667 items per 10,000 monthly credits and 10,000 items at most. When you switch the scenario back on, Make asks whether to process the old requests or delete them. Choose without thinking and you either replay a week of stale events or drop them. Make also deactivates a webhook that has not been connected to any scenario for more than 5 days, and the hook then returns 410 Gone.

Volume and timing: one item, slowly, versus many items, immediately

The test sent one item, and it sent it long after the record was created. Live runs reverse both.

On volume, Zapier holds tasks when a polling trigger returns more than 100 new records at once. The default limit is 100 and Company plans can raise it to 1,500 per Zap, and held tasks wait until someone decides to play them. Make returns a 429 error above 300 incoming webhook requests per 10 seconds. The first time most workflows meet these numbers is a bulk import, a list cleanup, or a backlog after an outage, which is also when nobody is watching. If a burst is the trigger, the guide to running an automation on existing records without emailing everyone covers the write side of that problem.

On timing, a test usually runs minutes or hours after you created the record. A live webhook fires within seconds. That gap matters whenever a later step searches for the record or something attached to it. A search that worked in the test can come back empty live, because many vendor search endpoints read an index that has not caught up yet. That is a specific failure with a specific fix, covered in why an automation cannot find the record it just created. Polling hides this for a while: Zapier polls every 15 minutes on the Free plan and every 1 to 2 minutes on paid plans, so upgrading a plan can expose a race that was always there.

State and error handling: what the editor quietly gave you

The last two rows are about what the editor provided that production does not.

n8n is the clearest example. Its documentation says production executions ignore all pinned data. If you pinned the output of an API call while building, every manual run reused that frozen response, and the live run is the first one that calls the API for real, with real credentials, real rate limits, and a real response shape. Pinning is a fine development aid. It is not evidence that the node works.

Error handling is where a failure becomes silent. n8n's Error Trigger documentation states that the Error Trigger only runs when an automatic workflow errors, so you cannot test an error workflow by running the workflow manually. If every test was manual, you have never seen your alert fire. Zapier records live errors in Zap History rather than the editor, and Make routes them through whatever error handlers you attached, which by default is none. In every case, the person who would have seen the failure in the editor is not looking.

This is the row that turns a five-minute fix into a three-week data problem. A failure you hear about the same day costs one bad record. A failure nobody hears about costs every record since launch. If your workflows do not yet tell you when they stop or fail, setting up an alert for an automation that stops running is the first thing to fix, before any of the other five.

How to test so the first live run is not the real test

You do not need a staging environment for this, although why a duplicated workflow is not one is worth reading before you build one. You need a test that exercises all six rows:

  1. Use three real records, not the sample. Pull them from run history or the source app: one complete, one with every optional field blank, and one with odd characters or formatting. If the workflow already ran live and failed, use that exact record.
  2. Enter through the production path. Publish the n8n workflow and hit the production URL. Turn the Zap on and create a brand-new record rather than editing an old one. Send the Make webhook to the live scenario.
  3. Send a burst. Five records at once is enough to show whether steps run in parallel, collide, or hit a per-second limit.
  4. Run it at live speed. Create the record and let the trigger fire immediately, with no pause for you to check it first.
  5. Force a failure. Break one step on purpose, for example with n8n's Stop and Error node, and confirm the alert reaches a person.

Then check the result in the same place you will check it in a month: Zap History, the Executions tab, the scenario log. If the run is not visible there, it will not be visible when it fails.

What to do with the workflow that just broke

Open the failed live run and the passing test side by side, and walk the table above row by row. The record row catches most cases in a few minutes. If the inputs match, check the trigger path next, because a workflow that never ran leaves no failed run to compare.

If you are fixing the same class of failure across many workflows, that is a sign the problem is the testing practice rather than any one flow. We build workflow automation systems with production-path tests and failure alerts from the first run, and if you want a second set of eyes on a workflow that keeps breaking, get in touch.

Frequently Asked Questions

The test used one sample record, usually a recent and complete one, pulled once on demand. The live Zap polls on a plan-based interval, skips any item whose ID it has already seen, and receives real records with blank or oddly formatted fields. Open Zap History, compare the input of a failed run with the test record, and the difference is almost always visible in one field.

Manual and production executions are different paths in n8n. Production executions ignore all pinned data, a webhook workflow only answers on its production URL once the workflow is published, and the Error Trigger does not run for manual executions. Check the Executions tab rather than the editor, because production data does not display on the canvas.

Polling triggers deduplicate by ID. Zapier's developer documentation says that when a Zap is first turned on, it stores the ID of every existing item, and it only runs for IDs it has not seen. An edited record keeps its ID, so a new-record trigger ignores it unless the trigger is built to combine the ID with an update timestamp.

A test sends one item. Zapier holds tasks when a polling trigger returns more than 100 new records at once (the default flood protection limit), and Make returns a 429 error above 300 incoming webhook requests per 10 seconds. A bulk import or a backlog after downtime is the usual first time a workflow sees that volume.

Replay three real records instead of the sample: a complete one, one with every optional field blank, and one with unusual characters or formatting. Then trigger the workflow the way it will actually run, through the production webhook URL or the live schedule, and confirm a failure reaches your error alert.

SOURCES & CITATIONS

  1. How Zap triggers work — Zapierhttps://help.zapier.com/hc/en-us/articles/8496244568589
  2. Choose your Zap's flood protection limit — Zapierhttps://help.zapier.com/hc/en-us/articles/14230420160269-Choose-your-Zaps-flood-protection-limit
  3. Manual, partial, and production executions — n8nhttps://docs.n8n.io/workflows/executions/manual-partial-and-production-executions/
  4. Webhooks — Makehttps://help.make.com/webhooks

About Alexey Yushkin

Alexey is the founder of GENERAL INFORMATICS LLC. He designs and ships AI and automation systems for businesses and operators across the US.

Connect on LinkedIn

Related reading

Want this kind of system in your business?

We build practical AI and automation systems for operators. Send us your current workflow and we will show you what to automate first.