It Works! Now
What?

Fast Iteration for AI Capabilities in Airflow

Alex Guglielmone Nemi

Engineering Lead @ Amazon

2026-09-01

Two Worlds

A Real Example

Which Bin Tomorrow?

? ? RECYCLING RUBBISH FOOD AUGUST 27 Thursday Rubbish Recycling Rubbish Recycling Rubbish 2 weeks 2 weeks

That's a whole month of rubbish in my garden

The Agent Figures It Out

* My setup included custom context and tools. Measure your own.

It Works. Now What?

Fast Iteration Needs a Cheap Loop

Before you delegate,

build the check that tells you it's right.

You could also call it executable intent.

The Smallest Check

class BinDayFacts(BaseModel):
    next_collection_date: date
    containers_due: set[str]


def test_ordinary_week():
    facts = ask_agent(recorded("2026-08-22"))   # the council's real response, saved

    assert facts == BinDayFacts(
        next_collection_date=date(2026, 8, 27),
        containers_due={"BLUE RECYCLING WHEELIE BIN", "FOOD BOX"},
    )
PASSED  test_ordinary_week
FAILED  test_collection_pulled_forward
        - containers_due={'BLACK RUBBISH WHEELIE BIN', 'FOOD BOX'}
        + containers_due=set()

This can be called an AI eval. It is a deterministic assertion on AI-generated output.

From Agent Skill to Airflow Dag

the new requirements, as tests on a plain function
def test_alerts_only_when_the_collection_is_tomorrow():
    assert decide(facts(next="2026-08-27"), today="2026-08-26") == "send"
    assert decide(facts(next="2026-08-27"), today="2026-08-22") == "explain"

def test_does_not_say_it_twice():
    assert decide(facts(next="2026-08-27"), today="2026-08-26",
                  already_alerted="2026-08-27") == "explain"

before  claude -p "Do I need to put the bins out? …" --model opus --allowedTools WebFetch WebSearch Bash(curl:*)

@dag(schedule="0 7 * * *", tags=["bins", "agent"])
def bin_day():
    SKILL = load_skill(".claude/skills/bin-day/SKILL.md")  # the same file, byte for byte
    agent = Agent(
        BedrockConverseModel(model, provider=BedrockProvider(...)),
        output_type=BinDay,         # the pydantic model the check compares
        instructions=SKILL.body,   # the markdown, loaded not retyped
        tools=[get_addresses, get_collections], # and nothing that can message me
    )
    @task
    def ask_agent() -> dict:
        today = get_current_context()["logical_date"].date()  # never the clock
        return agent.run_sync(today=today).output
    @task.branch(outlets=[ALERT_STATE])   # plain Python decides
    def decide(payload) -> str: ...
    ask_agent() >> decide(...) >> [send(...), explain(...)]

* Trimmed for width. Or skip the hand-rolling: apache-airflow-providers-common-ai ships AgentOperator and AgentSkillsToolset.

Change the Model

Change the Model

* One capability, one recorded world, three runs per model. Illustrative of the method, not an authoritative comparison of these models.

Making the Loop Faster

naive agent run check, by a judge model cheaper check agent run an assertion costs milliseconds independent cases that share no state can overlap isolated setup teardown isolation is not free

Move More Out of Inference

Agent Skill

Find the Loop

A goal worth pursuing Otherwise nobody pays the time, and nobody pays the tokens either.
the loop delegation change → run → check
Levers it can change The prompt, the model, the tools, the workflow, the runtime. Plus the permissions to change them.
Somewhere to experiment It has to reach everything that affects the outcome.
A way to compare A baseline, a check, and the evidence the check produces.

Define the Box

every lever so far make the box bigger the Dag the whole of Airflow scheduler · triggerer · workers real AWS Glue

Every lever so far has been inside a dag: the prompt, the model, the tools, the workflow.

Now my levers are: kill a worker · restart the triggerer · configure the schedulerCheap, because Airflow ships images. Just a compose file I can destroy.

Hypothesize and Prove

I'm not building this feature. I want to know it works for my cases.

What if the worker dies while the job is still running?
What if the job already finished?
What if the job id is gone?
And the same kill with the feature off, as a control.
a Glue job of 75 seconds, not the forty minutes that's normal
orphan detection cut from Airflow's 5 min default to ~40 seconds

Here the check is a hypothesis. True or false doesn't matter. Either way I know.

Simulate the Effect

the decision made by the agent what I am afraid of and how far it spreads something in between a fake notifier a read-only credential a dry run, a mock, a sandbox the real world money, messages, deletes

Two Kinds of Control

Can it do this at all? Permissions. Tool exposure. Credentials. Sandboxes. RBAC.
If it must never happen, make it impossible.
Did it do the right thing? Checks. Evals. Assertions. Recorded cases. Regression suites.
If it is allowed but needs judgement, evaluate it.

Evals are not security.

Traces: How The Answer Came To Be

run c542976d model claude-opus-5, effort high 22 turns 20 tool calls 6 refused 124.2s 989,639 tokens
  2.1s  user      do I need to put the bins out? postcode <redacted>
  6.2s  Skill     bin-day loaded
 12.8s  ToolSrch  select:WebFetch                                ok
 41.8s  Bash      fetch the council waste page                   ok
 50.5s  Bash      inspect form fields and endpoints             REFUSED
 58.1s  Bash      inspect form fields                           REFUSED
 62.4s  Bash      inspect form fields                           REFUSED
 71.6s  Bash      inspect form fields and endpoints              ok
 81.2s  Bash      look up addresses for postcode                 ok
 92.5s  Bash      extract UPRN for <redacted>                   REFUSED
101.7s  Bash      fetch collection schedule for UPRN             ok
111.4s  Bash      clean up temporary lookup files               REFUSED
115.8s  Bash      clean up temporary lookup files               REFUSED
124.2s  answer    Thursday 27 Aug · BLUE RECYCLING WHEELIE BIN, FOOD BOX

* Trimmed for width: 14 of 26 events. 16×Bash, 2×Read, 1×Skill, 1×ToolSearch. 989,639 tokens for three rows of a table — and ~30k of that was my harness, before the task started.

Prod Traces Capture Your Next Check

You were never going to predict every case. We didn't manage that with ordinary software either.

a run does something new you notice the trace has the whole case it becomes a check
tests/test_bin_day.py
  def test_ordinary_week(): ...
  def test_collection_pulled_forward(): ...
+ def test_bank_holiday_moves_the_whole_week():
+     facts = ask_agent(recorded("2026-08-25"))
+     assert facts.next_collection_date == date(2026, 8, 28)

Make Good Checks Easy to Write

unit tests test-driven, undogmatically integration tests system tests evals
/SkillBuilder shipped alongside Agent Skills. You just talk to it.
→
/EvalBuilder Encode good eval patterns one you know what "good" looks like

A qualifying exam for a domain, before we let an agent work in it.

Dag Operator Dag Code Generator Localizer … yours

Takeaways

1  Build executable checks for your loops So you can delegate strongly.
2  Find the right box A task, a Dag, the whole Airflow image. Or bigger.
3  You may not need the inference at all AI all around your system, and workflows with none in them.

Evals are not security. If something must never happen, prevent it rather than evaluating it.

Questions?

https://linktr.ee/alex.guglielmone.nemi

Writings · Github · LinkedIn

Alex Guglielmone Nemi · Airflow Summit 2026 · Austin, TX

Tux, after Larry Ewing