Runway Unveils GWM Worlds 2, an AI Model That Builds Interactive Video Worlds
The New York AI company says the tool generates continuous 720p video and audio that respond to typed commands in real time.
Runway, the AI video company, announced a new model called GWM Worlds 2 that it says generates continuous, interactive video and audio worlds that respond to a user's commands as they happen. The company described the system in a research post on its website.
According to Runway, the model produces 720p video at 24 frames per second along with audio at 48,000 Hz, and it keeps generating new footage as long as a user keeps interacting with it, rather than playing out a fixed clip. Runway said users first set up a scene, including its environment, characters, visual style, physical rules and ambient sound, and then steer that world with short text commands aimed at either a character or the scene itself, alongside continuous camera movement. A single command might tell a character to walk, deliver a line of dialogue, or change the weather, Runway said.
Typed Commands Instead Of a Script
Runway said the model supports both first-person and third-person views, including walking, driving, riding and flying, and that the camera can move independently of whatever a character is doing. Characters can also hold conversations, with users able to set voice, tone and language, and Runway said lip movement and delivery are generated to match the speech.
Runway calls its underlying command format a "WorldPrompt." It splits a scene into things that stay constant, such as layout, lighting and the rules of the world, and things that change over time, such as an individual action tied to a start and end time. The company said multiple actions can overlap, and that different users can each control a separate character, which it said could support multiplayer use.
Three Ways To Run It
Runway laid out three ways the model could be used. In the first, a person writes out a full sequence of timed actions in advance and the model generates the entire video from that script, which Runway said suits filmmaking. In the second, generation pauses at decision points for a user to choose what happens next, which the company said fits visual novels. In the third, the model generates continuously while a user's actions change the video in real time, which Runway said is the hardest version to pull off because it requires reacting within a fraction of a second, faster than a user could type a full prompt. For that mode, Runway said its demo binds keys and mouse clicks to preset prompts, though it said writing out fuller prompts ahead of time produces better results because they can describe more of the scene.
Runway called the model a follow-up to an earlier version, GWM Worlds, which it showed in December, with the new version adding generated audio and finer control over subjects and scenes. The company did not say when or whether the model will be released publicly, or what it might cost to use.