Build a Multi-Agent Sequential Pipeline with Semantic Kernel

September 28, 2026 • guides
Python

This guide is adapted from Microsoft Semantic Kernel's step2_sequential.py, licensed under the MIT Licence. Code blocks are reproduced exactly from the source; prose is original.


Orchestrating a single LLM call is straightforward. The interesting engineering challenge arrives when your task has genuine phases — analysis, composition, refinement — where each phase requires a different specialisation and where the output of one stage is the raw material for the next. Semantic Kernel's SequentialOrchestration makes that pipeline explicit and inspectable: you declare a list of agents, hand the runtime a single task string, and the framework routes each agent's output to the next agent's input without you writing any glue logic. The pattern maps cleanly onto anything with a natural editorial or analytical funnel: content pipelines, code-review chains, data-enrichment flows, or document summarisation ladders.

This guide is for engineers who already know how to make a single LLM call and want to move up to coordinated, multi-step pipelines. You do not need deep ML knowledge, but you do need to be comfortable with Python async, Azure credentials, and the idea that each agent is a prompt-plus-model configuration, not a fine-tuned model. For a sense of where the broader agent-orchestration landscape is heading, Google's open-source Kubernetes orchestrator is an instructive parallel — see our coverage of Google AX, the Kubernetes agent orchestrator — though Semantic Kernel runs entirely in-process and requires no cluster.

The sample calls Azure OpenAI three times per invocation, once per agent. No GPU is required; everything runs on a laptop or a CI runner, provided your Azure subscription has a deployed chat model endpoint.


Prerequisites

  • Python 3.11 or later
  • semantic-kernel package installed (pip install semantic-kernel)
  • azure-identity package installed (pip install azure-identity)
  • An Azure OpenAI deployment with a gpt-4o or equivalent chat model
  • Azure CLI authenticated in your shell (az login) — the code uses AzureCliCredential, so no API key is needed at the code level
  • The AZURE_OPENAI_ENDPOINT and AZURE_OPENAI_DEPLOYMENT_NAME environment variables set, or an AzureChatCompletion service configured to pick them up from your environment

No Docker, no special hardware. A standard developer laptop with an internet connection is sufficient.


Step 1: Define Your Specialised Agents

The first structural decision is how many agents you need and what each one is responsible for. This example uses three agents arranged in an editorial pipeline: one that extracts structured concepts from a product description, one that writes marketing copy from those concepts, and one that edits and polishes the draft. Each is a ChatCompletionAgent — a lightweight wrapper around a chat-completion service and a system-level instruction string.

The order of the list is the execution order. SequentialOrchestration is not inferring a dependency graph; it is following the list index.

# Copyright (c) Microsoft. All rights reserved.

import asyncio

from azure.identity import AzureCliCredential

from semantic_kernel.agents import Agent, ChatCompletionAgent, SequentialOrchestration
from semantic_kernel.agents.runtime import InProcessRuntime
from semantic_kernel.connectors.ai.open_ai import AzureChatCompletion
from semantic_kernel.contents import ChatMessageContent


def get_agents() -> list[Agent]:
    """Return a list of agents that will participate in the sequential orchestration.

    Feel free to add or remove agents.
    """
    credential = AzureCliCredential()

    concept_extractor_agent = ChatCompletionAgent(
        name="ConceptExtractorAgent",
        instructions=(
            "You are a marketing analyst. Given a product description, identify:\n"
            "- Key features\n"
            "- Target audience\n"
            "- Unique selling points\n\n"
        ),
        service=AzureChatCompletion(credential=credential),
    )
    writer_agent = ChatCompletionAgent(
        name="WriterAgent",
        instructions=(
            "You are a marketing copywriter. Given a block of text describing features, audience, and USPs, "
            "compose a compelling marketing copy (like a newsletter section) that highlights these points. "
            "Output should be short (around 150 words), output just the copy as a single text block."
        ),
        service=AzureChatCompletion(credential=credential),
    )
    format_proof_agent = ChatCompletionAgent(
        name="FormatProofAgent",
        instructions=(
            "You are an editor. Given the draft copy, correct grammar, improve clarity, ensure consistent tone, "
            "give format and make it polished. Output the final improved copy as a single text block."
        ),
        service=AzureChatCompletion(credential=credential),
    )

    # The order of the agents in the list will be the order in which they are executed
    return [concept_extractor_agent, writer_agent, format_proof_agent]

The instruction strings do double duty: they set each agent's persona and specify the expected output format. The WriterAgent instruction constrains output to roughly 150 words and a single text block. That constraint matters because the next agent in the chain receives exactly that text as its input — if an upstream agent produces structured JSON when the downstream agent expects prose, the pipeline degrades silently rather than failing loudly.


Step 2: Add an Observation Callback

Add the callback function that lets you see each agent's output as it is produced rather than waiting for the final result. This is your primary debugging surface: if one agent produces bad output, the callback tells you immediately which agent it was and what it said, without requiring you to instrument the orchestration itself.

def agent_response_callback(message: ChatMessageContent) -> None:
    """Observer function to print the messages from the agents."""
    print(f"# {message.name}\n{message.content}")

The ChatMessageContent object carries the agent's name alongside its content, so you can correlate output to the agent that produced it. You can extend this function to write to a log file, push to a metrics system, or trigger a webhook — any synchronous side effect is valid here.


Step 3: Assemble the Orchestration and Start the Runtime

With agents and callback defined, compose the SequentialOrchestration and start the InProcessRuntime. The runtime is the message-passing substrate that moves ChatMessageContent objects between agents; starting it before invoking the orchestration is a hard requirement.

async def main():
    """Main function to run the agents."""
    # 1. Create a sequential orchestration with multiple agents and an agent
    #    response callback to observe the output from each agent.
    agents = get_agents()
    sequential_orchestration = SequentialOrchestration(
        members=agents,
        agent_response_callback=agent_response_callback,
    )

    # 2. Create a runtime and start it
    runtime = InProcessRuntime()
    runtime.start()

InProcessRuntime runs inside the current Python process — there is no network hop, no sidecar service, no broker to configure. This makes local development fast, but if your process dies mid-pipeline there is no checkpoint to resume from. For production workloads that need durability, you would swap the runtime for a distributed one, though that is outside this tutorial's scope.


Step 4: Invoke, Collect Results, and Shut Down

The final segment invokes the orchestration with the task string, awaits the result with a timeout, prints it, and then gracefully stops the runtime. The timeout=20 argument to get() is a ceiling in seconds; if the full pipeline has not completed by then, the call raises an exception rather than hanging indefinitely.

    # 3. Invoke the orchestration with a task and the runtime
    orchestration_result = await sequential_orchestration.invoke(
        task="An eco-friendly stainless steel water bottle that keeps drinks cold for 24 hours",
        runtime=runtime,
    )

    # 4. Wait for the results
    value = await orchestration_result.get(timeout=20)
    print(f"***** Final Result *****\n{value}")

    # 5. Stop the runtime when idle
    await runtime.stop_when_idle()


if __name__ == "__main__":
    asyncio.run(main())

stop_when_idle() is cooperative: it waits until all in-flight messages have been processed before tearing down the runtime. Calling runtime.stop() instead would be an abrupt shutdown and could interrupt agents still generating tokens, producing incomplete or lost output.


Parameter Reference

Parameter Type What it controls When to change it
name str Agent identifier surfaced in ChatMessageContent.name and logs Always set a meaningful name — it is your only handle in the callback
instructions str System-level prompt defining persona, task scope, and output format Whenever output format or specialisation changes; be explicit about expected output shape
service ChatCompletionClientBase The underlying LLM endpoint; each agent can use a different model or deployment Use a cheaper, faster model for extraction agents and a higher-capability model for creative or editorial agents
members (on orchestration) list[Agent] Ordered list of agents; index determines execution sequence Reorder to change pipeline flow; add agents to add stages
agent_response_callback Callable Observer fired after each agent completes; receives full ChatMessageContent Omit only if you have no observability need; always include during development
timeout (on get()) float (seconds) Maximum wall time to wait for the full pipeline to complete Increase for longer pipelines or slower models; decrease to fail fast in latency-sensitive contexts

What to Watch Out For

Silent format drift between agents. The biggest practical failure mode is one agent producing output in a shape the next agent was not designed to consume. The orchestration passes text through mechanically; it does not validate schema or structure. If ConceptExtractorAgent starts outputting JSON because the model decided that was cleaner, WriterAgent may quote raw JSON in its copy. Fix this by being explicit in instructions about output format, and use the callback to catch drift early.

Timeout sizing. Size timeout to (number_of_agents × expected_max_agent_latency) + buffer. A pipeline that exceeds the timeout raises rather than returning partial results, and you will not know how far it got unless your callback was logging.

Single-point credential. All three agents share one AzureCliCredential instance. If your Azure session expires mid-run, every agent fails. For service deployments, use a managed identity or a service principal with a token that outlives your longest expected pipeline run.

No retry logic. SequentialOrchestration does not automatically retry a failed agent. A transient HTTP 429 from Azure OpenAI will propagate as an exception, terminating the pipeline at the failing stage. Wrap your main() invocation in retry logic at the caller level, or configure the AzureChatCompletion service with its built-in retry policy before passing it to the agent.

InProcessRuntime is not durable. A process crash loses all in-flight work. For workloads where the pipeline takes minutes and the inputs are expensive to regenerate, checkpoint the callback output to persistent storage at each stage — this gives you a manual resume path even without a distributed runtime.

Token accumulation across stages. Each agent receives the previous agent's full output as its input message. In a long pipeline, the cumulative context passed forward can grow large. If an upstream agent produces verbose output and your model has a tight context window, downstream agents may truncate or behave unexpectedly. Explicitly constrain output length in each agent's instructions, as the source's WriterAgent does with the "around 150 words" guidance.

If your inputs come from user-supplied text rather than internal product descriptions, prompt injection is a meaningful attack surface — a topic worth reviewing in light of recent agent security disclosures.


Where to Go Next

Once the sequential pattern is running, the natural next step is conditional routing — pipelines where the output of one agent determines which agent runs next, rather than a fixed sequence. Semantic Kernel's AgentGroupChat and planned orchestration primitives cover this. You should also look at how to swap InProcessRuntime for a distributed runtime when you need durability, and consider how to integrate tool calling so that individual agents can query APIs or databases rather than working purely from context. The Semantic Kernel repository ships additional step files in the same directory as step2_sequential.py that walk through those patterns in the same style as this guide.

Frequently asked questions

What Python version does Semantic Kernel SequentialOrchestration require?

The guide and the upstream source both require Python 3.11 or later. Earlier versions lack the async features that InProcessRuntime and the orchestration primitives rely on.

Can each agent in SequentialOrchestration use a different model or deployment?

Yes. Each ChatCompletionAgent accepts its own `service` argument, so you can pass a different AzureChatCompletion instance — pointing at a different deployment or even a different endpoint — to each agent. A common pattern is using a cheaper, faster model for extraction stages and a higher-capability model for creative or editorial stages.

What happens if an agent fails or the timeout is exceeded in SequentialOrchestration?

A transient error such as an HTTP 429 from Azure OpenAI propagates as an exception and terminates the pipeline at the failing stage. If the full pipeline does not complete within the seconds passed to `orchestration_result.get(timeout=...)`, the call raises rather than returning partial results. The callback output is your only record of how far the pipeline progressed before failure, so logging it to persistent storage is advisable in production.

Is InProcessRuntime suitable for production workloads?

InProcessRuntime is suitable for development, testing, and short-lived jobs where durability is not required. Because it runs inside the current Python process, a crash loses all in-flight work with no checkpoint to resume from. For pipelines that run for minutes or that process expensive inputs, checkpoint the callback output to persistent storage at each stage, or swap to a distributed runtime when one becomes available in Semantic Kernel.

Do I need an API key to run this example?

No. The sample uses `AzureCliCredential` from the `azure-identity` package, which picks up the token from an active `az login` session. You do need the `AZURE_OPENAI_ENDPOINT` and `AZURE_OPENAI_DEPLOYMENT_NAME` environment variables set, and your Azure account must have access to a deployed chat model endpoint.

How do I prevent output format drift between agents in a sequential pipeline?

Be explicit in each agent's `instructions` string about the expected output format and length — the source's WriterAgent, for example, specifies 'around 150 words' and 'a single text block'. Use the `agent_response_callback` to inspect each agent's raw output during development. If an upstream agent starts producing JSON when the downstream agent expects prose, the orchestration passes it through mechanically without validation, and the pipeline degrades silently rather than raising an error.

Related Guides