HeyGen ships Avatar V to keep AI clones recognizable across longer videos HeyGen released Avatar V on September 17, a video-reference-conditioned avatar model that generates videos from a 15-second recording and is priced at 48 credits per minute for a video look, available in HeyGen Studio and Video Agent. HeyGen's own technical report, comparing Avatar V against four other video-generation models on a 70-case cross-scene test set, claims it led on identity preservation, lip synchronization and generation quality, though the company notes some competitors' outputs did not use matching speech audio. The release extends HeyGen co-founder and CEO Joshua Xu's bet on replacing the camera, and HeyGen's help documentation says Avatar IV remains the choice for photo-based looks and virtual or non-human characters. HeyGen ships Avatar V to keep AI clones recognizable across longer videos The model uses a 15-second recording to recreate a person's appearance and movement; HeyGen's claims of leading realism rely partly on its own testing. By Ryan Merket https://runtimewire.com/author/ryan-merket ยท Published Primary source: HeyGen https://www.heygen.com/blog/announcing-avatar-v Why it matters Avatar V shifts HeyGen's pitch from generating a talking likeness to maintaining a recognizable person across repeated business videos. The commercial test is whether viewers accept those generated performances as credible communication, not just whether the model wins its maker's benchmark. HeyGen https://runtimewire.com/models/fal/heygen-avatar4-digital-twin 's Avatar V generates videos from a 15-second recording, with the company pitching it as a way to keep a person's likeness consistent across scenes, outfits and longer scripts. The product announcement was dated September 17th, not October 10th, the date attached to an image URL that has since circulated for the story. For Joshua Xu @joshua xu https://x.com/joshua xu , HeyGen's co-founder and chief executive, the release continues a bet rooted in his earlier work on cameras. Xu studied computer science at Carnegie Mellon and worked at Snap on advertising systems and computational photography. He has described the founding idea as replacing the camera: using AI to make video without requiring people to film themselves. The new model takes that premise from making a face speak toward reproducing how a particular person moves and performs. Identity is the product claim Avatar V learns from a short video reference, which HeyGen says captures gestures, expressions and mannerisms. Users can then generate a presenter in different settings and clothing without recording every version. The product is aimed at tasks such as training, product demonstrations and customer communications, where a company might want to update a video without booking another shoot. That focus builds on work HeyGen has discussed before. In August, RuntimeWire reported on TAVR, HeyGen's video-reference system https://runtimewire.com/article/heygen-tavr-video-reference-talking-avatars , which used as many as 48 reference frames to preserve a person's identity across generated scenes. Avatar V applies the same identity-consistency problem to HeyGen's commercial avatar product, using the short reference clip as an input for repeat video generation. Avatar V is designed for real people and video-based looks, while HeyGen's help documentation says Avatar IV remains the choice for photo-based looks and virtual or non-human characters. Avatar V is available in HeyGen Studio and Video Agent. The published rate is 48 credits per minute for a video look, a cost teams producing content at volume will need to account for. A benchmark with a company behind it HeyGen calls Avatar V the world's most realistic avatar model. That superlative remains a company claim. A technical report from HeyGen Research compares the system with four other video-generation models using a 70-case cross-scene test set. The report says Avatar V led on identity preservation, lip synchronization and generation quality, but it is HeyGen's own evaluation, not an independent audit. Its researchers also note that some competitors' outputs did not use matching speech audio, so those comparisons assessed visual quality separately. The report describes Avatar V as a video-reference-conditioned system, meaning it uses the reference footage directly to model identity rather than relying only on a fixed identity representation. That design targets the problem the product announcement emphasizes: keeping a generated person recognizable while the scene or camera angle changes. A small test set built and assessed by the model's maker cannot settle how well the system performs across the full range of real-world footage or use cases. Xu said in June that HeyGen had surpassed $200 million in annual recurring revenue, with more than 30 million users and adoption at 85% of Fortune 100 companies. Those are company-reported figures, not audited results. In 2024, Benchmark led HeyGen's $60 million Series A at a $500 million valuation, with participation from Conviction, Thrive Capital and Bond Capital, according to Bloomberg. Avatar V is a product bet aimed at getting those business users to use synthetic presenters for more than quick internal explainers. A convincing digital stand-in could make routine video production cheaper and easier to update. For a business relying on a familiar employee or executive to deliver the message, however, the model must preserve qualities beyond facial likeness and make the generated performance credible enough that viewers accept it as that person's communication. HeyGen's launch argues that Avatar V clears that bar. Its own benchmark supports a narrower claim about measured identity and video quality; whether customers treat the output as a substitute for recording remains the more consequential test.