BestAI Newsroom research note

This evergreen history article uses authoritative archives and official records. Exact dates are used when documented; gradual inventions and rollouts are described as periods rather than being assigned a misleading single birthday.

Quick facts

  • HeyGen was founded in 2020 by Joshua Xu and Wayne Liang and initially operated under the name Movio.
  • The product lets users create presenter-style videos from scripts without filming a person for every version.
  • Its avatar, voice-cloning and lip-synchronization systems became popular for marketing, training and localization.
  • Video Translate became a signature feature by preserving a speaker’s appearance while changing spoken language.
  • By 2026, HeyGen had developed into a broader platform for avatars, agents, translation and personalized video.

Origins and founding

HeyGen was created to reduce the time and cost of traditional video production. Recording a presenter normally requires a camera, location, lighting, editing and repeated takes. Producing the same message in many languages multiplies that work. Joshua Xu and Wayne Liang believed synthetic presenters could make routine business video as easy to update as a slide deck.

The company began in 2020 and first became known as Movio. Early customers used template-based avatars to turn scripts into short promotional or educational clips. The product later adopted the HeyGen name as it expanded beyond a narrow spokesperson-video tool.

The product takes shape

HeyGen combined stock digital presenters with custom avatars created from recordings of a real person. A user could type a script, select a voice and produce a talking video. Improvements in facial animation, lip movement and neural speech made results more convincing and reduced the mechanical appearance common in older avatar systems.

Video Translate attracted wider attention because it changed the spoken language of an existing clip while attempting to preserve the speaker’s voice, timing and mouth movement. This was useful for creators and companies that wanted one recording to reach audiences in many countries.

Technology and major features

The platform uses speech synthesis, voice conversion, facial animation, tracking and generative image or video models. Custom avatars learn visual and motion patterns from approved footage. Translation systems transcribe the original speech, translate its meaning, generate a matching voice and retime facial movement.

HeyGen later added interactive avatars and agent-like experiences, allowing a digital person to respond in real time rather than only reading a fixed script. APIs and enterprise tools support automated personalization, where names, languages or offers can change for many recipients while the basic video structure remains consistent.

Growth and wider influence

HeyGen became popular among sales teams, educators, marketers, online creators and internal communications departments. It lowered the barrier to producing presenter videos and made localization possible for organizations that could not hire actors and studios in every market.

The service also helped normalize synthetic humans in ordinary business communication. Instead of viewing avatars only as entertainment characters, companies used them for onboarding, product explanation and customer engagement. Competition with Synthesia, D-ID and other platforms accelerated quality improvements.

Challenges, criticism and responsibility

Realistic avatars create serious consent and impersonation risks. A person’s face or voice should not be cloned without clear authorization. Translated speech can also falsely imply that someone personally spoke words they never reviewed. Disclosure is especially important in politics, finance, medicine and news.

HeyGen uses verification and moderation processes, but users remain responsible for lawful source material and honest presentation. Another limitation is emotional nuance: an avatar can communicate clearly while still missing the subtle timing and spontaneous expression of a human performance.

Where it stands in 2026

By 2026, HeyGen was positioned as an end-to-end synthetic-video platform that included avatars, translation, voices, personalization and interactive agents. Enterprise customers valued rapid localization and consistent branding, while independent creators used the service to produce content without a studio.

Future development is likely to improve real-time conversation, full-body movement, scene generation and cultural adaptation. The technology’s acceptance will depend on reliable consent records and transparent labeling so that useful localization does not become effortless deception.

Timeline

YearLocationEventWhy it mattered
2020Los Angeles, United StatesThe company is foundedBegins building AI-presenter video tools.
2021–2022Global web marketMovio avatar platform gains usersMakes script-to-presenter video accessible through templates.
2022–2023GlobalThe product adopts the HeyGen brandSignals expansion beyond its original avatar-video identity.
2023–2024GlobalVideo Translate and improved custom avatars launchMakes multilingual localization a defining feature.
2025–2026Enterprise and creator marketsInteractive avatars, agents and APIs expandTurns HeyGen into a broader synthetic communication platform.

Frequently asked questions

Who founded HeyGen?

HeyGen was founded by Joshua Xu and Wayne Liang.

What was HeyGen previously called?

The company’s product was previously known as Movio.

What is HeyGen Video Translate?

It translates spoken video into another language while attempting to preserve the speaker’s voice, timing and synchronized mouth movement.

Can anyone clone another person in HeyGen?

Legitimate use requires authorization and compliance with verification, consent and platform rules; cloning someone without permission can violate rights and laws.

Final perspective

HeyGen shows how generative video can solve a practical problem: communicating the same idea through many presenters and languages. Its benefits are strongest when the speaker has consented and the synthetic nature of the content is clear. Without those safeguards, the same realism can weaken trust in recorded speech.

Sources and references