Deploy an LTX-2 Two-Stage Text-to-Video API on Modal H200 GPUs

October 9, 2026 • guides
Serverless

Deploying heavy generative media models to production is traditionally an exercise in resource frustration. Teams are usually forced into a painful compromise: either provision expensive GPU instances that sit idle during off-peak hours burning through infrastructure budgets, or risk devastating cold-start latencies that ruin the end-user experience. Unlike specialised architectural use-cases—such as the spatial reasoning engines we covered in the Reka Rho article—a robust text-to-video pipeline requires orchestrated multi-stage diffusion, synchronised audio processing, and heavy text encoding.

This guide demonstrates how platform engineers and machine learning practitioners can build a production-grade inference service on serverless infrastructure. We will construct a two-stage video generation pipeline using LTX-2, a 19-billion parameter model. By offloading lifecycle management to Modal, the service scales dynamically based on incoming requests, meaning you only pay for the compute seconds used during generation.

Be prepared for the hardware constraints involved. Generating 1024x1536 resolution video demands immense VRAM; attempting this workflow on mid-tier hardware will trigger out-of-memory crashes during the spatial upscaling phase. The upstream example targets Nvidia H200 accelerators. The source does not publish a canonical per-generation time or cost; exact numbers depend on current H200 availability, queue depth, and prompt length. The decisions below focus on making the run possible at all rather than on a particular bill.

Note: This guide is adapted from Modal Examples' ltx2_two_stage.py, available at https://github.com/modal-labs/modal-examples under the permissive MIT licence.

Prerequisites

To replicate this deployment, you need an active Modal account and the Modal CLI installed and authenticated on your local machine. You must also have a Hugging Face account to access the Gemma-3 12B instruction-tuned weights used by the text encoder. Visit the Gemma-3 repository on Hugging Face to accept the user agreement, then generate an access token. Finally, store this token in Modal's secret manager under the name huggingface-secret with the key HF_TOKEN.

Step 1: Build the container environment

Serverless workloads require deterministic, pre-compiled environments. Building a container image dynamically allows us to inject precise versions of deep learning libraries and system-level utilities before the application even boots.

import time
from pathlib import Path

import modal

# We pin an LTX-2 commit and install its three subpackages.
# Torch is installed, along with a Flash-Attention 3 (as a wheel)
# for faster inference on GPUs with the Hopper
# [SM architecture](https://modal.com/gpu-glossary/device-hardware/streaming-multiprocessor-architecture).

ltx2_commit = "28c3c73"

image = (
    modal.Image.from_registry("nvidia/cuda:12.6.1-devel-ubuntu24.04", add_python="3.12")
    .apt_install("git", "ffmpeg")
    .uv_pip_install(
        "torch==2.7.0",
        "torchaudio==2.7.0",
        "transformers>=4.52,<5",
        f"git+https://github.com/Lightricks/LTX-2.git@{ltx2_commit}#subdirectory=packages/ltx-core",
        f"git+https://github.com/Lightricks/LTX-2.git@{ltx2_commit}#subdirectory=packages/ltx-pipelines",
        f"git+https://github.com/Lightricks/LTX-2.git@{ltx2_commit}#subdirectory=packages/ltx-trainer",
        "https://huggingface.co/alexnasa/flash-attn-3/resolve/main/128/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl",
        extra_index_url="https://download.pytorch.org/whl/cu128",
        extra_options="--index-strategy unsafe-best-match",
    )
    .env(
        {
            "HF_XET_HIGH_PERFORMANCE": "1",
            "PYTORCH_ALLOC_CONF": "expandable_segments:True",
        }
    )
    .entrypoint([])
)

We start from an Nvidia CUDA 12.6.1 development image, ensuring compatibility with the latest PyTorch distributions. Notice that we pin a specific Git commit for the LTX-2 repositories. Research repositories evolve aggressively; tracking the main branch directly introduces severe deployment risks when upstream authors introduce breaking API changes.

We also explicitly install a precompiled wheel for Flash Attention 3. This is crucial for Hopper architecture cards such as the H200 to ensure the attention mechanism achieves maximum throughput. The environment variable PYTORCH_ALLOC_CONF set to expandable_segments:True acts as a safeguard against VRAM fragmentation, allowing PyTorch to efficiently manage the massive memory footprint of a 19B parameter model without hitting artificial allocation limits.

Step 2: Attach persistent volumes and application state

If a serverless container must download over 50 gigabytes of model weights from Hugging Face upon every single invocation, the latency will be catastrophic. We solve this by attaching network volumes that persist data independently of the container lifecycle.

# ## Volumes

# Model weights are cached to a Volume at HuggingFace's default cache path.
# Generated videos are saved to a separate output Volume.

model_volume = modal.Volume.from_name("ltx2-models", create_if_missing=True)
output_volume = modal.Volume.from_name("ltx2-outputs", create_if_missing=True)

OUTPUT_DIR = Path("/output-videos")

with image.imports():
    import torch
    from huggingface_hub import hf_hub_download, snapshot_download
    from ltx_core.loader import LTXV_LORA_COMFY_RENAMING_MAP, LoraPathStrengthAndSDOps
    from ltx_core.model.video_vae import TilingConfig, get_video_chunks_number
    from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline
    from ltx_pipelines.utils.constants import (
        DEFAULT_AUDIO_GUIDER_PARAMS,
        DEFAULT_NEGATIVE_PROMPT,
        DEFAULT_VIDEO_GUIDER_PARAMS,
    )
    from ltx_pipelines.utils.media_io import encode_video

app = modal.App(
    "example-ltx2-two-stage",
    image=image,
    volumes={
        "/root/.cache/huggingface": model_volume,
        OUTPUT_DIR: output_volume,
    },
    secrets=[modal.Secret.from_name("huggingface-secret")],
)

Here, we declare two distinct volumes: one dedicated to caching model weights and another strictly for output persistence. By mapping model_volume to the default /root/.cache/huggingface directory within the container, any internal library calling standard Hugging Face download functions will transparently read from our high-speed network cache instead of the public internet.

The with image.imports(): context manager ensures that the heavy Python imports are evaluated within the remote container environment rather than on your local machine, preventing local dependency conflicts.

Step 3: Initialise the pipeline on an H200

We must now define the core class responsible for managing the hardware allocations and loading the weights into memory. This setup phase runs exactly once when a container spins up, preparing the application to accept multiple subsequent generation requests.

# ## Inference

NUM_FRAMES = 121  # ~5s at 24 fps
FRAME_RATE = 24
WIDTH = 1536
HEIGHT = 1024


@app.cls(gpu="H200", timeout=30 * 60, scaledown_window=15 * 60)
class LTX2TwoStage:
    @modal.enter()
    def setup(self):
        """Download model weights and initialize the two-stage pipeline."""
        torch.set_float32_matmul_precision("high")

        repo = "Lightricks/LTX-2"
        checkpoint_path = hf_hub_download(repo, "ltx-2-19b-dev.safetensors")
        upsampler_path = hf_hub_download(
            repo, "ltx-2-spatial-upscaler-x2-1.0.safetensors"
        )
        distilled_lora_path = hf_hub_download(
            repo, "ltx-2-19b-distilled-lora-384.safetensors"
        )
        gemma_dir = snapshot_download("google/gemma-3-12b-it-qat-q4_0-unquantized")
        model_volume.commit()

        distilled_lora = [
            LoraPathStrengthAndSDOps(
                distilled_lora_path, 1.0, LTXV_LORA_COMFY_RENAMING_MAP
            )
        ]

        self.tiling_config = TilingConfig.default()
        self.pipeline = TI2VidTwoStagesPipeline(
            checkpoint_path=checkpoint_path,
            distilled_lora=distilled_lora,
            spatial_upsampler_path=upsampler_path,
            gemma_root=gemma_dir,
            loras=[],
        )

The decorator @app.cls(gpu="H200", ...) explicitly requests a massive compute footprint. The scaledown_window argument is a crucial economic control: it dictates that once a container has finished a generation request, it will wait in an active state for 15 minutes before terminating. If your service receives intermittent traffic, this window keeps the weights warm in VRAM, avoiding the setup penalty for subsequent requests.

Inside the setup method, we pull the base checkpoint, the spatial upscaler, the distilled LoRA, and the Gemma-3 weights. The immediate model_volume.commit() call forces these downloaded files to sync back to the persistent volume, ensuring subsequent boot sequences bypass the network download entirely. We also instantiate a TilingConfig; managing VAE decoding via spatial tiling is non-negotiable for high-resolution video frames, as processing everything at once would reliably exhaust even the H200's capacity.

Step 4: Generate and encode the video

Once the models are securely loaded into VRAM, we construct the method that will handle the incoming prompt, orchestrate the inference, and multiplex the generated audio and visual arrays into a usable MP4 format. The following block continues the LTX2TwoStage class from Step 3.

    @modal.method()
    def generate(self, prompt: str) -> None:
        """Generate a video from a text prompt and save it to the output Volume."""
        print(f"Generating {NUM_FRAMES} frames ({NUM_FRAMES / FRAME_RATE:.0f}s) ...")
        print(f"Prompt: {prompt}")
        start = time.time()

        with torch.no_grad():
            video, audio = self.pipeline(
                prompt=prompt,
                negative_prompt=DEFAULT_NEGATIVE_PROMPT,
                seed=42,
                height=HEIGHT,
                width=WIDTH,
                num_frames=NUM_FRAMES,
                frame_rate=FRAME_RATE,
                num_inference_steps=40,
                video_guider_params=DEFAULT_VIDEO_GUIDER_PARAMS,
                audio_guider_params=DEFAULT_AUDIO_GUIDER_PARAMS,
                images=[],
                tiling_config=self.tiling_config,
                enhance_prompt=True,
            )
            print(f"Generated in {time.time() - start:.0f}s")

            safe = "".join(c if c.isalnum() or c == " " else "-" for c in prompt)
            filename = f"{int(time.time())}_{safe[:80].strip().replace(' ', '_')}.mp4"
            output_path = OUTPUT_DIR / filename

            encode_video(
                video=video,
                fps=FRAME_RATE,
                audio=audio,
                audio_sample_rate=24_000,
                output_path=str(output_path),
                video_chunks_number=get_video_chunks_number(
                    NUM_FRAMES, self.tiling_config
                ),
            )
        output_volume.commit()
        print(f"Saved to Volume `ltx2-outputs` at {filename}")

We wrap the entire execution in torch.no_grad() because tracking gradients during forward-pass inference wastes critical memory and compute cycles. We pass 40 inference steps to the pipeline; this is generally the sweet spot for distilled models, balancing visual fidelity against generation latency.

The encode_video function offloads the heavy lifting to ffmpeg, which we installed in Step 1. It takes the raw tensor outputs and interleaves the 24,000 Hz audio track with the 24 FPS video frames. Crucially, the resulting file is written into OUTPUT_DIR. Just as we did with the weights, we must explicitly call output_volume.commit() to flush the file from the container's temporary filesystem back to the persistent network drive, ensuring it survives when the serverless container eventually scales down.

Step 5: Submit from the CLI

To test the deployment, we define a local entrypoint. This function runs on your local workstation, triggers the remote infrastructure, and blocks until the video is encoded on the cloud.

@app.local_entrypoint()
def main(
    prompt: str = "A cathedral made of ice, northern lights dancing overhead, camera slowly pushing forward through the nave",
):
    LTX2TwoStage().generate.remote(prompt=prompt)

By calling .remote(prompt=prompt), Modal intercepts the invocation, provisions the necessary H200 container if one is not already warm, executes the class initialisation, and then passes the string payload directly to the generate method.

You can submit the same concrete prompt with the CLI:

modal run ltx2_two_stage.py --prompt "A cathedral made of ice, northern lights overhead"

After the remote call completes, run modal volume ls ltx2-outputs to list the generated MP4, then use modal volume get ltx2-outputs with the timestamped MP4 name returned by the listing.

GPU selection parameters

Selecting the right hardware architecture for generative video heavily dictates what limits you will encounter. The table below illustrates how different serverless configurations affect complex generative deployments.

GPU Model VRAM Architecture Suitability for LTX-2 (19B)
A10G 24 GB Ampere Insufficient VRAM. Will fail during model load.
A100 80 GB Ampere Functional, but requires strict memory management and slower attention implementation.
H100 80 GB Hopper Excellent throughput, fully supports Flash Attention 3, but spatial upscaling VRAM limits can be tight.
H200 141 GB Hopper Optimal. Vast VRAM overhead entirely prevents VAE decoding out-of-memory errors at 1024x1536.

What to watch out for

When adapting this pipeline for a production application, several failure modes and hidden costs require your attention.

VRAM Spikes During VAE Decoding: The diffusion steps themselves are surprisingly manageable; the true danger lies in the final Variational Autoencoder decoding step, where latent representations are projected back into high-resolution pixel space. If you alter the TilingConfig or attempt to push the resolution beyond 1024x1536, VRAM consumption will spike sharply. Even on an H200, improper tiling configurations will result in instant out-of-memory terminations.

Hugging Face Rate Limiting: The reliance on hf_hub_download and snapshot_download during container setup introduces a fragile dependency on an external network. If your model_volume fails to cache correctly and your service scales up to dozens of concurrent instances to handle a traffic spike, you risk hitting rate limits or silent connection timeouts from Hugging Face. Always monitor the execution time of the setup block; if a warm volume still takes seconds to initialise, your cache path is likely misconfigured.

Scale-down Economics: The scaledown_window=15 * 60 parameter instructs Modal to keep the H200 instance alive for fifteen minutes after a request. While this prevents a cold-start penalty, it also means you pay for fifteen minutes of idle time after the final user drops off. If you are wrapping this class in a synchronous web endpoint, you must profile your real-world traffic patterns and adjust this window. A shorter window saves money but drastically degrades the experience for sparse, infrequent users.

Next steps

Now that you have a functioning remote generation class, the logical next step is adapting it for a frontend integration. You can modify the script by importing Modal's web-endpoint decorators, such as @modal.web_endpoint, to wrap the generate function inside a FastAPI route. That allows external applications to submit prompts via standard POST requests and retrieve presigned URLs for the final MP4 files.

Frequently asked questions

What GPU do I need to run the LTX-2 two-stage pipeline?

The example pins an Nvidia H200 with 141 GB of VRAM. LTX-2 is a 19B-parameter diffusion model with a text encoder and spatial upscaler, and the 1024x1536 VAE decode is the largest memory consumer, so H200 gives enough headroom to avoid tiling-related out-of-memory errors. The script can technically target other GPUs, but A10G-class cards are not viable.

Why does the guide use two Modal volumes?

The `ltx2-models` volume caches downloaded Hugging Face weights at `/root/.cache/huggingface`, so containers do not re-download LTX-2 and Gemma checkpoints on every cold start. The `ltx2-outputs` volume stores encoded MP4 files in `/output-videos` after `output_volume.commit()` flushes them from the container. This separates model cache from generated media and keeps both persistent after scale-down.

How do I retrieve a generated video from Modal?

Run `modal volume ls ltx2-outputs` to list files, then use `modal volume get ltx2-outputs` with the timestamped MP4 name returned by the listing. The `generate` method writes a timestamped MP4 to the output volume. You need the Modal CLI installed and authenticated to run these commands.

Do I need a Hugging Face token for LTX-2?

Yes, the text encoder uses Google's Gemma 3 12B instruction-tuned weights, which require accepting a license on Hugging Face. Create a token and store it as a Modal Secret named `huggingface-secret` with the key `HF_TOKEN`. The app passes that secret into every container, so `snapshot_download` can access the Gemma repository.

What does `scaledown_window=15 * 60` do?

It keeps an idle Modal container warm for 15 minutes after the last request before scaling it down. This avoids the full setup penalty for successive requests but continues billing for idle H200 time for that window. Lower it for sparse traffic, raise it for bursty traffic.

Why does the code install Flash Attention 3 from a wheel?

LTX-2 runs on Hopper-generation GPUs such as the H200, and the wheel provides Flash Attention 3 for faster attention computation. The source also sets `PYTORCH_ALLOC_CONF=expandable_segments:True` to reduce memory fragmentation during long video diffusion runs.

Related Guides