Running Elastic Workflows: What the Execution Detail Tells You
In-stack automation as YAML. What actually happens when a workflow fires, and what four months of execution history says about the parts that hold up.
Elastic Workflows (Kibana 9.x) is automation that lives inside the stack. You write a workflow in YAML, give it a trigger, and it runs steps against Elasticsearch, Cases, connectors, HTTP endpoints and AI agents. No external orchestrator, no separate runtime, no credentials shuttled between systems.
We have sixteen of them on one SOC cluster. This post walks through three — an alert-noise detector, an ingestion monitor and a Fleet agent check — reading their real execution detail, and then checks what those runs imply against the 1,309 executions already in history.

(The full list runs past one screen; the capture above is the same table scrolled through and joined, so all sixteen rows are readable.)
The list view earns a second look before anything else. The Triggers and Steps column renders one icon per trigger and one per step type, so you can read a workflow’s shape — scheduled, five queries, an email — without opening it. The Enabled column is the number that matters most here: every one of the sixteen is off. More on that at the end.
The anatomy
A workflow is four blocks: name, triggers, steps, enabled. Here is the simplest one we run, in full:
name: Duplicate Alert Grouping
triggers:
- type: scheduled
with:
every: "10m"
steps:
- name: grouped_alerts
type: elasticsearch.esql.query
with:
format: json
query: |
FROM .alerts-security.alerts-*
| WHERE @timestamp > NOW() - 10 minutes
| STATS duplicate_count = COUNT(*)
BY host.name,
user.name,
source.ip,
kibana.alert.rule.name
| WHERE duplicate_count >= 5
| SORT duplicate_count DESC
- name: suppression_summary
type: console
with:
message: |
Duplicate Alert Groups Detected
{{ steps.grouped_alerts.output.values | json }}
enabled: false
Thirty lines, and it’s an alert-noise detector. The mechanic to internalise is the last one: steps.<name>.output is how data moves. Every step publishes its result under its own name and any later step templates it in. That single convention is most of what you need.
Templating is Liquid-flavoured:
| Expression | Use |
|---|---|
{{ steps.grouped_alerts.output.values }} |
a previous step’s result |
{{ consts.some_value }} |
a workflow constant |
{{ event.alerts[0].kibana.alert.rule.name }} |
the triggering alert |
\| default: "N/A" |
survive a missing field |
\| size |
count rows without adding a step |
\| json |
dump a result into a message body |
| default: earns its keep fastest. Alert documents are not uniformly populated, and a workflow that renders {{ event.alerts[0].user.name }} into a case title produces blank titles the moment an alert has no user.
Three trigger types, and the choice decides how you can test:
scheduled— polling and reporting. Runs onevery: "10m"/"15m"/"1h"/"24h".alert— fires from a detection rule, with the alert at{{ event.alerts[0].* }}. You cannot invoke this by hand.manual— analyst-invoked, and the only trigger you can drive from the editor on demand.
Across the sixteen: scheduled ×8, alert ×5, manual ×3. Two thirds are scheduled or manual — the unglamorous kind that checks data is still arriving and agents are still checked in.
Duplicate Alert Grouping
One query, one console step. Nothing leaves the cluster, so it is safe to invoke against production data from the editor.

Success in 836ms. The execution pane is the reason to run rather than read. Three things appear that YAML cannot tell you:
- A per-step timeline with real durations —
grouped_alerts44ms,suppression_summary2ms. The two steps account for 46ms of an 836ms run, so roughly 95% of the wall-clock is orchestration, not work. On a 10-minute schedule that is irrelevant; it matters the moment you reach for sub-minute intervals. - The trigger that fired, shown as its own node above the steps. Invoking a
scheduledworkflow by hand is recorded asmanualand flagged a test run — history keeps the two apart. - Every step’s input and output, individually inspectable.
That last one is where the learning is. Click a step and you get its output as a field table:

And the same payload as raw JSON:

{
"took": 28,
"is_partial": false,
"completion_time_in_millis": 1789990198343,
"documents_found": 1,
"values_loaded": 3,
"rows_emitted": 8,
"bytes_read": 0,
"read_nanos": 0,
"cpu_nanos": 5240190,
"columns": [ { "name": "duplicate_count", "type": "long" }, … ]
}
Read that carefully, because it is the shape of an ES|QL step’s output and not what most people assume when they write {{ steps.x.output }}. The rows live under .values, their schema under .columns. Alongside them:
is_partial— the field to care about. An ES|QL query that timed out or hit a shard failure returnsis_partial: truewith a successful step status. A downstream step that files a case off partial results reports a partial picture as a complete one, and nothing anywhere is marked failed. If a workflow acts on query results, gate it onis_partial.documents_found: 1/values_loaded: 3/rows_emitted: 8— three different counts, none interchangeable: documents scanned, values materialised, rows returned. Onlyrows_emittedis the size of the result your next step will template over.took: 28vs the step’s own 44ms — Elasticsearch spent 28ms; the rest is the step’s round trip.columns— name and type per column, which is what lets you parse| jsonoutput downstream instead of guessing at positional arrays.
Confirm the output shape before you act on it. Build every workflow with a console step dumping | json, check the structure, then wire in the real action. Guessing the shape and going straight to cases.createCase is how you generate a dozen malformed cases.
Data Ingestion Monitor
The second workflow is the shape most SOC automation actually takes: query, then deliver. It runs every: "15m", looks for log data streams that have gone quiet, and emails the analysts when any of them exceed the delay threshold.

This is a scheduled execution, not a test invocation — triggeredBy: scheduled, isTestRun: false — which is why the trigger node reads scheduled rather than manual. Success in 2s: ingestion_check 537ms, send_ingestion_alert 1s, alert_console 81ms.
Two things are worth pulling out of those numbers. The query is the cheap part: 537ms of a 2-second run. And the email step costs roughly twice what the query does, which is the ratio that holds across the whole library.
The step detail is where an email step stops being a black box:

envelopeTime : 137
messageTime : 294
envelope.from: soc-alerts@example.com
envelope.to : soc-alerts@example.com
rejected : -
response : 250 2.0.0 OK <message-id@example.com>
accepted[0] : soc-alerts@example.com
ehlo[] : SIZE, PIPELINING, DSN, ENHANCEDSTATUSCODES,
AUTH LOGIN XOAUTH2, 8BITMIME, BINARYMIME, CHUNKING
You get the raw SMTP conversation: envelope timing, the server’s capability list, the accepted and rejected recipient arrays, and the literal 250 2.0.0 OK. That rejected array is the field to build on — a step can report success while the mail server silently drops a recipient, and rejected is the only place that shows.
This workflow is also where the Liquid control flow lives. Its message body iterates the query result and branches on the delay magnitude:
{%- for source in steps.ingestion_check.output.values %}
Dataset:
{{ source[1] }}
Last Event Received:
{{ source[4] | date: "%Y-%m-%d %I:%M:%S %p" }} PKT
Ingestion Delay:
{%- if ... %}
Note source[1] and source[4] — positional indexing into the row array. It works, and it is brittle: add a column to the EVAL and every index below it shifts, silently, in the body of an email nobody reads closely. This is the argument for reading .columns rather than counting positions.
The query that was correct when it was written
Look at the status bar in both captures: 1 error. The workflow ran clean in May; the same YAML no longer plans today.
Cannot use field [data_stream.dataset] due to ambiguities being mapped as
[2] incompatible types: [keyword] in [227 indices] and [text] in [1 index]
One newly created data stream mapped data_stream.dataset as text while 227 others map it keyword, and ES|QL refuses to aggregate across a field with two types. Nothing about the workflow changed — the index set underneath it did.
A FROM logs-* wildcard is a promise that every index it matches agrees on your field types, and that promise is not yours to keep. Cast it (data_stream.dataset::keyword) or scope the pattern. It is the clearest example in our library of why a disabled workflow is not a safe workflow: this one would have started failing on a schedule, into an inbox, without anyone touching it.
Fleet Agent Offline Check
The third runs every: "1h" against Fleet’s own agent index and reports agents that are enrolled but not checking in.

Success in 1s — offline_agents 124ms, send_agent_alert 1s. Same split as before: the query is a rounding error, the delivery is the run.
The useful observation is what the step can reach. FROM .fleet-agents works from a workflow, and so does FROM .kibana-event-log-* and FROM .alerts-security.alerts-*. Workflows query dot-prefixed system indices directly, which is how the interesting operational workflows get built: rule health from the event log, agent health from the Fleet index, alert noise from the alerts index. None of it needs an exporter or an external collector.
Three workflows, three different index families, every query under 550ms. The elasticsearch.esql.query step is the load-bearing part of this feature — 32 of the 79 steps in our library — and it behaved identically in each.
What 1,309 executions say
Every run is a document in .workflows-executions, which makes execution history queryable rather than just browsable. The editor shows it per workflow:

Today’s run sits at the top, above months of scheduled runs at 637ms, 854ms, 669ms, 508ms and one 4s outlier. Consistency like that is the argument for reading history instead of a single run: one 4s outlier among sub-second runs is cluster load, not a code change.
Across the whole index:
| Executions | 1,309 (12 May – 21 Sep 2026) |
| Completed | 1,216 |
| Failed | 92 |
| Cancelled | 1 |
| Success rate | 92.9% |
| Trigger split | scheduled 1,112 / manual 197 |
| Duration | min 228ms · avg 2,453ms · max 254,488ms |
That 254-second maximum against a 2.5-second average is the distribution to notice. It is not a slow query. Grouping the library by what a workflow actually touches:
| Workflow shape | Runs | Failed | Avg duration |
|---|---|---|---|
Query-only (ES|QL + console) |
1,004 | 0 | ~800ms |
| Query + case / connector delivery | 198 | 41 | 1.1–2.6s |
| Query + LLM agent | 19 | 12 | 15,960ms |
The pattern is unambiguous. Workflows that only query Elasticsearch approach 100% reliability at sub-second latency. Workflows that call out — connectors, third-party HTTP, LLM agents — own every failure and every slow run. The LLM-agent workflows average 16 seconds and fail more often than they succeed; the 254-second outlier is one of their runs. The pure-query workflows ran over a thousand times without a single failure.
The 92 failures
| Error type | Count |
|---|---|
ConnectorExecutionError |
55 |
ResponseError |
21 |
Error |
9 |
InputValidationError |
4 |
CaseError |
3 |
60% are connector failures — the workflow logic ran, the delivery didn’t. Grouped by message:
| Cause | Count |
|---|---|
[401] Unauthorized (three distinct endpoints) |
18 |
[403] Forbidden: @font-face{font-family:Poppins… |
15 |
verification_exception (ES|QL) |
14 |
[ENOTFOUND] getaddrinfo … your-threat-feed.example.com |
5 |
parsing_exception (ES|QL) |
4 |
Three are worth pulling out.
The @font-face one is the best error in the corpus. A 403 whose body is CSS means the HTTP step was handed an HTML login page instead of JSON. The endpoint wasn’t down and the request wasn’t malformed — a session had expired and something returned a sign-in page. Fifteen executions failed against a login screen. An http step will happily POST at anything; it cannot tell you that what came back was a web page.
your-threat-feed.example.com is a placeholder that reached production. Five scheduled executions failed on DNS because a template value was never replaced. It failed loudly, which is the good outcome — the same oversight in a WHERE clause fails silently and returns nothing.
18 ES|QL exceptions. verification_exception is a field that doesn’t exist or has become ambiguously mapped; parsing_exception is a syntax error. Both are caught at query planning, so the step fails fast and cheap. Most are preventable by running the query in Discover first — but not all, as the ingestion monitor shows: a query can be correct when written and fail months later because the indices underneath it changed.
Failure detail is positional, and that is the point
The execution pane marks each step individually, not just the run. In the recurring failure shape here, several ES|QL steps finish green and only the final delivery step goes red, carrying:
error.type : ConnectorExecutionError
error.message : [401] Unauthorized: The authentication credentials are not valid.
That tells you the gathered result sets are intact and the only broken thing is a credential in a connector. A single top-level “failed” status would have sent someone re-reading ES|QL for an hour.
It also exposes a design consequence: there is no partial-success state and no resumption. A workflow like that does all of its work and then discards it because a token expired at the last step. The next scheduled run redoes every query from scratch. If the gathering phase is expensive, put it behind its own workflow and pass results through an index rather than a step boundary.
Three things the execution detail changed our minds about
Control flow isn’t missing — it’s in the wrong place to be reviewable
Not one of the sixteen workflows has a condition or if key at step level. Branching exists, but it lives inside step bodies as Liquid: three workflows use {% for %} loops and three use {% if %} blocks, all of it inside message and body strings, as the ingestion monitor shows.
So conditional logic in this system is embedded in the templates that render output. It works. But a branch inside a string is a branch you cannot see from the step timeline, cannot time separately, and cannot inspect in the execution pane — the pane shows the rendered result, not which arm produced it. Every branch is a path that executes only during an incident, which is the worst time to discover it was wrong.
Push decisions into ES|QL instead. Filter and aggregate in the query so the step below receives exactly the rows it should act on. The query language is more expressive than the orchestrator, its errors are caught at planning time, and its output is what the execution pane shows you.
consts are not secrets
consts are stored in the workflow YAML, and that YAML is returned verbatim by GET /api/workflows. Anything you put there is readable by anyone who can call that endpoint, and lands in every export, backup and screenshot of the workflow.
We found a live third-party API token sitting in a consts block. It reads as harmless while you are authoring, because it looks like configuration.
Credentials belong in connectors, referenced by connector-id / agent-id — which is exactly how the email step above is wired: connector-id: soc-alerts-notify, no password anywhere in the YAML. The same goes for recipient addresses and internal hostnames. A workflow YAML is a shared, API-readable object; treat everything in it as published.
enabled: false is not a neutral resting state
All sixteen are disabled. It is genuinely easy to build these — a useful workflow is thirty lines — and that is the trap. A library of disabled automation is a library of untested assumptions.
The three workflows here make the case. Each one ran successfully, and reading the detail surfaced something the YAML did not: an is_partial flag nobody was checking, a rejected recipient array nobody was checking, positional source[4] indexing that breaks silently when a column is added, and a wildcard aggregation that stopped planning because one new index mapped a field differently.
Either a workflow runs and earns its keep, or it is a draft that will be quietly wrong by the time someone needs it.
What we'd tell someone starting
Write the ES|QL first. 32 of 79 steps are queries, and 18 of 92 failures are query errors. Get the query right in Discover, then wrap a workflow around it.
Cast or scope every wildcard field. A FROM logs-* aggregation is only as stable as the mapping agreement across every index it matches, and that agreement breaks without warning.
Confirm the output shape. Rows are at .values, their schema at .columns, the row count at rows_emitted. Gate anything consequential on is_partial, and read rejected on delivery steps.
Index by name, not by position. source[4] is a time bomb in an email body.
Default every templated field. | default: on anything from an alert document or a STATS ... BY.
Any polling workflow that writes needs an idempotency guard. Tag what you acted on; exclude that tag from your own search. Without it, a five-minute poller re-files the same ticket forever.
Keep the expensive gathering separate from the fragile delivery. There is no partial success and no resume. A 401 at the last step throws away everything the earlier steps did.
Secrets, recipients and hostnames go in connectors, never consts.
Workflows will not replace a general-purpose orchestrator and does not try to. What it removes is the glue: no service to operate, no API client to maintain, no credentials shuttled between systems, and your automation is queryable YAML sitting next to the data it acts on. Our own numbers say to lean into exactly that — the pure-query workflows ran over a thousand times without a single failure. Everything that reached outside the cluster is where the 92 failures are.
Screenshots are from a live Kibana 9.x cluster, captured at 1920px. Execution figures are from .workflows-executions as of 21 September 2026. Recipient addresses, operator names, execution URLs, internal hostnames and a third-party ticketing product name were rewritten to neutral placeholders in the browser before each capture; every figure, timing, step name, query and error string is unmodified.