The post-trained MiniMax H3 variant pairs fal's model research with its serving stack to generate five-second videos in roughly three seconds.
By RuntimeWire Staff · Published
Primary source: PR Newswire
Why it matters #
Gur and Yurtseven are moving fal up the stack from hosting other labs' models to modifying weights and serving systems together. H3 Max tests whether model-specific optimization can give a generative-media platform an edge over broader inference providers and better-funded video labs.
Burkay Gur and Gorkem Yurtseven, the co-founders of San Francisco-based fal, released H3 Max on September 1, pushing their generative media platform beyond hosting other labs' models and deeper into developing the systems that determine how those models perform.
H3 Max is post-trained from the open-weights MiniMax H3 video model. fal Research added training data intended to improve prompt adherence and visual quality, while fal's inference engineers optimized the serving stack alongside the evolving model. In a September 1 release, fal said the result can generate a five-second video in roughly three seconds.
According to fal's technical announcement, H3 Max delivers about 35 times the throughput of MiniMax's official H3 endpoint. fal also compared H3 Max with 12 leading video models in internal testing. The results remain a fal-reported comparison rather than a standardized independent speed benchmark.
The launch fits the original bet Gur and Yurtseven made when they founded fal in 2021. Gur previously led machine-learning development at Coinbase after working at Oracle, while Yurtseven had been a software developer at Amazon. The longtime friends explored infrastructure ideas during the pandemic before experiments with Stable Diffusion narrowed their focus to generative media. Gur told TechCrunch in 2024 that they expected AI-generated media to reshape much of the content people consume.
H3 Max turns that thesis into a more ambitious product strategy. fal is applying its infrastructure work during model development instead of waiting for a finished set of weights to arrive from another lab.
The inference engine becomes part of the model
fal's technical announcement describes a joint research and systems process. Its researchers evaluated post-training checkpoints through head-to-head human preference studies covering overall quality, prompt understanding and aesthetics. At the same time, the inference group tested changes to precision, sampling and execution, retaining optimizations only when fal's evaluations showed that output quality held up.
The approach matters for fal's business because model APIs are easy to compare when providers expose the same weights. Latency, throughput, uptime and price become the practical points of competition. By modifying the model and designing the runtime around it, fal can offer an endpoint that is harder for a broader GPU platform to reproduce by the same public weights.
fal competes for inference workloads with providers including Replicate, Modal, RunPod and CoreWeave, according to an industry analysis of fal. Those businesses span hosted model APIs, serverless compute and GPU clouds. fal's narrower pitch centers on generative media and optimization for individual models, with model weights, sampling configuration and the serving engine tuned together. H3 Max is the clearest test yet of whether that specialization produces a defensible difference.
Competition is also intensifying among video-model developers. Runway raised $315 million at a reported $5.3 billion valuation in February 2026. Kuaishou's Kling 3.0 supports multimodal input and output, native audio and clips up to 15 seconds, while Luma's Ray3.14 emphasizes native 1080p and HDR output. In that field, fal is trying to compete through the combination of model quality, serving speed and API distribution.
The work also moves fal closer to the model labs whose products it distributes. Through its generative media platform, fal gives developers one API surface for models built by outside labs, including MiniMax. H3 Max gives fal a direct claim over post-training and serving performance, even though MiniMax supplied the base model.
H3 Max is available through separate text-to-video and image-to-video endpoints, while fal says it retains MiniMax H3's multimodal context and synchronized audio and video capabilities.
Benchmark leadership depends on the category
fal's claim that H3 Max ranks first on independent benchmarks applies to specific image-to-video categories rather than the broader text-to-video leaderboard.
fal reported first-place results for H3 Max in image-to-video categories on Design Arena and Artificial Analysis. On the current Artificial Analysis text-to-video leaderboard, however, H3 Max ranked third at 1,235 Elo, behind Alibaba's Wan 3.0 and Google's Gemini Omni Flash.
The leaderboard listed H3 Max at $2.40 per minute, compared with $12 for Wan 3.0 and $6 for Gemini Omni Flash. Those figures are a dated benchmark snapshot rather than fal's current endpoint price. H3 Max can lead an image-to-video category while ranking third in the broader text-to-video comparison.
fal also ran its internal preference study against 12 leading video models. The company says H3 Max placed first across overall quality, prompt understanding and aesthetics and won most head-to-head matchups. Those results use prompts, evaluators and methodology selected by fal.
Speed is the distribution strategy
For Gur and Yurtseven, speed has remained the organizing principle from fal's first Stable Diffusion experiments through H3 Max. Faster generation reduces the cost of producing multiple candidates, makes creative interfaces feel responsive and gives applications room to regenerate a failed result without turning each interaction into a long wait. fal's dated disclosures show how quickly that pitch gained traction. In its July 2025 Series C announcement, fal said its platform supported more than two million developers, tens of thousands of applications and more than 300 enterprise customers. fal also said revenue had increased 60 times over the preceding 12 months. Those figures predate H3 Max and are not current usage totals for the model.
fal has raised substantial capital to build around that demand. Bloomberg reported that its December 2025 Series D brought in $140 million and valued fal at $4.5 billion. In fal's Series D announcement, fal named Sequoia, Kleiner Perkins and NVIDIA as new investors and said its team had grown to 70 people. Across its disclosed rounds, fal has raised $337 million.
fal had previously announced a $125 million Series C on July 31, 2025. The first-party announcement names Meritech as the lead investor and Salesforce Ventures, Shopify Ventures and Google AI Futures Fund as new participants.
H3 Max shows where some of that capital is going: fal is building research capacity that can change model behavior, then using its infrastructure to turn those changes into a production endpoint. Gur and Yurtseven are betting that application developers will place heavy weight on how quickly, reliably and cheaply a model runs, regardless of which lab published the original weights.
The five-seconds-in-three result gives fal a clean demonstration. The harder test will be whether co-developing models and inference systems preserves that advantage as rivals improve their models, prices change and the video leaderboards reshuffle.