Developers Innovation

When Ads Are Shaped by the Scene: Exploring Generative AI

TL;DR

  • Generative AI enables scene-aware advertising by dynamically adapting ad creatives to match the visual and contextual elements of the content being streamed
  • Scene-level analysis enhances contextual relevance, allowing ads to align with mood, objects, environments, or themes within a specific moment of a video
  • This approach moves beyond traditional targeting, shifting from demographic and behavioral data toward real-time content-driven personalization
  • AI-driven ad adaptation improves viewer experience by making ads feel more native and less disruptive to the content flow
  • The convergence of video intelligence and generative AI unlocks new monetization models, enabling scalable creative variation while maintaining brand consistency and performance goals

At Bitmovin, we believe that the best products often start as “bold ideas” born out of curiosity, especially when exploring how new technologies like generative AI could shape video workflows. That is why we host quarterly two-day hackathons, events designed to give our engineers the freedom to step away from their daily roadmaps and experiment with cutting-edge technology. These 48-hour sprints are a safe space for creativity, where the goal isn’t immediate perfection, but rather learning, collaborating, and having fun with the “what ifs” of video technology.

While many of these projects remain interesting experiments, others, like our AI Contextual Advertising solution,have evolved from rough hackathon sketches into full-fledged product features.

In this blog, we take a closer look at AI My Ads (AIMA), an internal project that emerged from a recent hackathon and asks a simple but provocative question. What if video ads could be generated on the fly?

The foundation of AIMA: AI Scene Analysis

To understand AIMA, you first have to understand the engine that drives it. At Bitmovin, that capability is AI Scene Analysis. Instead of treating a video as a single block of content, it looks at video scene by scene, building context around what is happening in each moment. That shift in how video is understood is what opened the door to exploring new questions around advertising.

By analyzing video, audio, and text together, AI Scene Analysis builds a detailed picture of what is happening within each scene, from setting and activity to objects, characters, and overall tone. This level of insight is already used to support contextual advertising workflows, helping ensure ads are scheduled at moments that align with the content being watched.(Contextual Advertising).

The Spark: Closing the Loop

With that foundation in place, the team started asking a different question. The scene-level context they now had access to was rich and detailed, but it was only being applied to decisions around timing. Nothing in the workflow used that same understanding to influence the creative itself.

That limitation became the starting point for AIMA. Rather than trying to build a new ad system, the team focused on a narrower idea. Could generative AI take scene-level context and use it as an input for creating ad content that reflects the content it appears alongside?

Building the Prototype: A 48-Hour Sprint

Supported by our partners at Google, hackathon teams got access to tools that might not yet be part of our daily stack, such as Vertex AI and advanced Gemini models. For AIMA, the team embraced the “move fast and break things” mentality, utilizing a powerful mix of generative AI tools to build the prototype:

  • Coding with an AI Copilot (Claude): Used to speed up development and keep the focus on overall logic rather than syntax.
  • The Brains: We used Gemini 2.5 Flash to analyze the metadata provided by AISA. It generated specific prompts for video creation based on the context of the scene.
  • The Visuals: Using Veo 3.1 Preview was used to generate short video ad segments, with defined start and end frames to help the ads blend more naturally into playback. 
  • The Voice: To polish the experience, we integrated ElevenLabs to generate voiceovers, ensuring the ad sounded as professional as it looked.

The Fun Factor: Seeing it in Action

A hackathon isn’t a hackathon without a sense of fun. To demonstrate the capabilities of AI My Ads, the team didn’t want to just generate generic commercials, instead, ads were inserted in the well known movie, Pulp Fiction (for the laugh factor of course). Take a look at some screenshots below… are you able to spot the Ad?

*The generated ads shown here are entirely fictional and have no association with the film or its creators. Scenes are used solely to demonstrate the prototype.

Screenshot from AI generated video of Jules holding a shoe

Screenshot from AI generated video of Jules holding playing with building blocks

Screenshot from AI generated video of Vincent holding a rubber duck

Image is worth a hundred words they say, and what is better than an image, a video. Watch the following AI My Ads output we made including a video of our CEO, Stefan Lederer and be prepared to be surprised with his latest new product, Bit Schnitzel.

Jokes aside, these images and videos show how generative AI can be used to translate scene-level context into creative output.

From Prototype to Potential Product

AIMA is currently a prototype, offering a glimpse into what can be explored when scene-level video understanding is combined with generative AI. It was built to test ideas rather than deliver a finished product, and to better understand how creative workflows might evolve when context extends beyond ad placement alone.

Whether or not AIMA ever becomes a product is not the point. The value lies in the questions it raises and the conversations it starts, about how advertising, creativity, and AI might intersect in the future. As an experiment, it reflects the kind of exploration that helps teams learn faster, think differently, and challenge assumptions before they ever reach a roadmap.

For Bitmovin employees, you can explore “the wonderful world of AI generated ads” yourself via our internal demo links. For everyone else, stay tuned.


FAQs

How does generative AI improve contextual advertising?

Generative AI enables real-time modification of creative assets based on scene data making ads more relevant to the exact moment within the video rather than relying solely on audience demographics or cookies.

How is scene-aware advertising different from traditional ad targeting?

Traditional targeting relies heavily on user data (behavioral, demographic, or interest-based signals). Scene-aware advertising leverages video content intelligence, adapting creatives to what is happening on screen, thereby increasing contextual alignment and reducing reliance on personal data.

How does scene-level video intelligence enable generative advertising?

Scene-level analysis identifies semantic signals within video content (e.g., environment, action, mood). These signals inform generative AI systems, which adapt creative elements in alignment with the detected context.

Leto Baxevanaki

Senior Engineering Manager - Player Web SDK

Leto Baxevanaki is a senior engineering manager at Bitmovin, leading the Web Player SDK and TA engineering teams. Her main responsibilities include leading multiple engineering teams, aligning their efforts with product roadmap and continuously enhancing software quality and delivery speed. Leto is deeply interested in promoting a culture of innovation, encouraging experimentation and new ideas, and exploring new technologies to achieve business advantages.

Giuseppe Samela

Giuseppe Samela

Senior Engineer | Player

Giuseppe Samela is part of the Bitmovin Player Engineering Team and is working on the Web SDK. His focus is on improving the overall stability of the product, and ensuring the best QoE for their users can be achieved. He has more than 7 years of experience in the Video Streaming field, with insights from both the Industry and the Research world.

Mukul Kumar

Video Player Engineer

Wolfram Hofmeister

Senior Software Engineer | Player Web

Wolfram is a video streaming enthusiast with many years of experience in the industry and a Masters degree in Applied Informatics, specializing in Distributed Multimedia Systems. While his main expertise is in building video streaming players for the web, he's also passionate about video encoding and VR/AR applications.


Related Posts

Developers

Live Translation for Broadcast, Multi-Audio, Multi-Subtitle, and DVB-Subs With Bitmovin’s Live Encoder

Developers

Hackathon spotlight: Upgrading MoQ support in Bitmovin’s Player Web X

Join our newsletter and stay informed.

Get to hear first when we publish something new.