AI-generated video has been advancing at an eye-popping pace over the past 10 months, and Google's remarkable new spatiotemporal diffusion model, Lumiere, has changed the goalposts yet again. Lumiere can create very realistic or high-quality surreal video clips up to 5 seconds long. It can also animate static images or portions of images based on natural language text prompts to let you know what you want to see.
It can take a picture, clone the style of that picture, and then use that style to create a slew of videos on other topics that look and feel so similar they could have been produced by a branding agency.
It can use your own source video to turn everything into Lego, origami, or flowers - you just tell it.
As you can see from the demo above, Lumiere has the most advanced in-video feature we've seen to date. All you have to do is paint in the parts of the image you don't like, and Lumiere will automatically fill in that area with a beautiful effect that you might not even notice if you're not looking carefully. Ex-boyfriend shows up in your favorite video? It won't be long.
The relevant research team stated that Lumiere's "spatio-temporal U-shaped network architecture" can construct the entire length of the video at once - while previous models usually generate the start frame and end frame first, and then guess what will happen in the middle.
No matter how you do it, the results speak for themselves—this is the new state of the art in generative AI video.
For now, this is just a research project - so that Google doesn't have to heavily emasculate the system for copyright, disinformation, safety, hate speech, nudity, privacy, and various other policies - a process that will inevitably lead to a decrease in the quality of the output of these generative models.