<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:cc="http://cyber.law.harvard.edu/rss/creativeCommonsRssModule.html">
    <channel>
        <title><![CDATA[JioHotstar - Medium]]></title>
        <description><![CDATA[Product and Engineering notes from JioHotstar, India&#39;s leading OTT service. Want to work with us? Head on over to https://www.jiostar.com - Medium]]></description>
        <link>https://blog.hotstar.com?source=rss----dbc3fcbc7f07---4</link>
        <image>
            <url>https://cdn-images-1.medium.com/proxy/1*TGH72Nnw24QL3iV9IOm4VA.png</url>
            <title>JioHotstar - Medium</title>
            <link>https://blog.hotstar.com?source=rss----dbc3fcbc7f07---4</link>
        </image>
        <generator>Medium</generator>
        <lastBuildDate>Fri, 11 Sep 2026 07:57:10 GMT</lastBuildDate>
        <atom:link href="https://blog.hotstar.com/feed" rel="self" type="application/rss+xml"/>
        <webMaster><![CDATA[yourfriends@medium.com]]></webMaster>
        <atom:link href="http://medium.superfeedr.com" rel="hub"/>
        <item>
            <title><![CDATA[Journey of an Ad Request: The Hidden Engineering]]></title>
            <link>https://blog.hotstar.com/journey-of-an-ad-request-the-hidden-engineering-e7b7f9921e46?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/e7b7f9921e46</guid>
            <category><![CDATA[advertising]]></category>
            <category><![CDATA[adtech]]></category>
            <category><![CDATA[ads]]></category>
            <dc:creator><![CDATA[Ayush Kumar]]></dc:creator>
            <pubDate>Mon, 20 Jul 2026 10:22:24 GMT</pubDate>
            <atom:updated>2026-07-20T10:22:23.508Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y7C7f95213Hu_VW4b1sGfQ.png" /></figure><p>Advertising is a key source of revenue in most businesses. It certainly forms a key pillar for JioHotstar. There is so much technology hidden under the surface. In this blog, we will cover some of the complexity that gets unlocked when we render an ad on our content. Every 30-second ad break is a high-stakes performance involving millions of calculations, real-time bidding, and surgical precision.</p><p>When you open JioHotstar to watch a live sports event or simply a movie, you see a <strong>Fence*</strong> ad (static display) or a <strong>Preroll*</strong> or a <strong>Midroll* Ad</strong> (video). What you don’t see is the “Ad Decision” engine firing at millisecond speeds to decide: <em>Why this ad? Why now? And why you?</em></p><h3>Speak Like an Adtech Pro</h3><p>Getting your foot inside the door, we need to understand some jargon!</p><ul><li><strong>CPM (Cost Per <em>Mille</em>):</strong> The cost an advertiser pays for every <strong>1,000</strong> impressions.</li><li><strong>CPC</strong> — Cost per Click</li><li><strong>SSAI (Server-Side Ad Insertion):</strong> Ads are “stitched” into the video stream on the server. At JioHotstar scale, this often happens at a <strong>Cohort</strong> level where we are grouping similar users to a balanced personalization of Ads.</li><li><strong>CSAI (Client Side Ad Insertion)</strong>: The frontend client app(player) interrupts the video stream based on cue markers to show an ad. These cue markers are known well in advance when the user plays any content.</li><li><strong>SGAI (Server-Guided Ad Insertion):</strong> The server “guides” the player on what to do, but the player performs the final insertion. This allows for 1:1 personalization at a live event scale where we do not have the cue markers in advance.</li><li><strong>Programmatic vs. Direct:</strong> <strong>Direct</strong> is a handshake deal (e.g., “I want 1M views on an Ad”). <strong>Programmatic</strong> is an automated auction where we ask a 3rd-party SSP (Supply Side Platform) for an ad in real-time.</li></ul><h4>Some common Ad Formats on JioHotstar</h4><p><strong>Preroll Ad</strong>: The video ad you see before the actual content video starts.</p><p><strong>Midroll Ad: </strong>The video ad you see while in between the content.</p><p><strong>Fence Ad</strong>: The static display ad you see below the player on the app.</p><h3>The Demand Puzzle</h3><p>Imagine you’re at your local market where you see two establishments side-by-side:</p><ol><li><strong>A global smartphone brand</strong> launching its latest flagship wants <strong>10 million impressions across India</strong> to drive nationwide awareness.</li><li><strong>A Local Boutique:</strong> Wants<strong> 1000 clicks </strong>(CPC) specifically from users in Indiranagar, Bengaluru.</li></ol><p>Both want to run ads, but their goals, scale, targeting strategy, and optimization signals are fundamentally different. One is optimizing for reach[how many different types of customers saw it] and visibility[how many unique geographies was this seen in] at national scale. The other is optimizing for precision, locality, and more measurable performance. Our system must juggle thousands of these conflicting demands simultaneously.</p><p>This leads us to the ultimate engineering challenge: <strong>The Selection.</strong></p><h3>Zeroing In: The Ad Decision Logic</h3><p>How do we pick 2–3 ads out of thousands for a 30-second ad “pod”? At JioHotstar, we use a very intricate implementation of waterfall tiers approach mixed with our <a href="https://arxiv.org/pdf/1905.10928">PID</a> and <a href="https://research.google/pubs/shale-an-efficient-algorithm-for-allocation-of-guaranteed-display-advertising-2/">SHALE</a> pacing algorithms. These calculations are all performed and delivered in under 100ms, even during massive concurrency spikes like the final over of an IPL match.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dDkhihTyv3PR-c3qT7qy2A.png" /><figcaption>A sample of the throughput spikes on the adserver that we get during a cricket match</figcaption></figure><p>During the Live sports events, the ads services get a very spiky traffic giving little to no time for the services to react to organic scale-up if we are not prepared in advance. The peaks are observed during ad breaks while the API traffic becomes more predictable and lower during the innings break(20:40 hrs — 21:10 hrs) as we see it in the graph above.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*RwpaMjRfke3USxS92pbw3A.png" /><figcaption><em>Flowchart indicating the steps taken in the ad server for decisioning</em></figcaption></figure><ol><li><strong>The Relevance</strong></li></ol><p>Out of 10,000 potential candidate ads, which specific ads are actually eligible for <em>a</em> female user, on <em>a premium Android phone</em> device, watching a <em>comedy</em> movie, in <em>Indiranagar</em>, Bengaluru?</p><p>Every request involves a realtime filtering of ads across geography, user affinity, device signals and advertiser constraints, etc. Within a few milliseconds, the system eliminates everything that doesn’t qualify and what remains is a smaller candidate pool.</p><p><strong>2. The Experience</strong></p><p>Nobody wants to watch the same brand ad again and again till they learn it by heart! This is why the ad server then must check for the Frequency caps (FCaps) across the combination of millions of users and thousands of ads via some distributed cache.</p><p><strong>3. The Ad Pacing and Time Interval Capping</strong></p><p>During a single IPL match where millions are watching it live, how do we stop the ad delivery when it hits its 1M impression target? How do we ensure we do not deliver all the 10M ad impressions within the first few days of a month-long campaign which was intended to create a prolonged reach on the masses?</p><p>This is where the math gets heavy. We use pacing algorithms (like PID or SHALE) to tackle these problems and ensure a uniform distribution of ad impressions throughout the campaign period. This also helps in preventing Over-delivery(OD) and Under-delivery(UD) of ads.</p><p><strong>3. The Priority Decisioning</strong></p><p>Out of 10 Ads that are qualifying for that one spot, which one brings out the most value or is the most important to be selected?</p><p>We resolve this through a <strong>Tiered Waterfall </strong>approach, where selection then proceeds tier by tier.</p><ul><li>Higher priority campaigns are evaluated first.</li><li>Lower tiers are considered only if the ad pod is not filled completely.</li><li>Any leftover duration is filled with in-house promos.</li></ul><p>If the selected ad is a direct deal, we immediately return the VAST response. But if it’s a programmatic ad, then we depend on third party SSP demand where fill rate uncertainty (guaranteed vs open exchange) becomes a factor.</p><h3>The Programmatic Challenge: Fill Rate</h3><p>Once we’ve ranked our eligible ads, the final step is fulfillment. If the winner is a Direct Deal Ad, we serve it immediately. But if it’s a Programmatic Ad, we’re essentially asking an external auction house (SSP): “<em>Hey, do you have an ad for this user right now?”</em></p><p>This is where the Fill Rate comes in. It’s the probability of actually getting an ad back. This becomes the make-or-break metric. Think of it as a reservation vs. a walk-in at a busy restaurant.</p><p><strong>Programmatic Guaranteed <em>(The Reservation)</em>:</strong> We’ve pre-negotiated this. The advertiser has committed to the volume of ads, so the Fill Rate is usually 80% to 100%.</p><p><strong>Open Exchange <em>(The Walk-in)</em></strong>: We throw the request out to the highest bidder in the open market. Because it depends on real-time competition and prices, the <em>Fill Rate</em> can plummet to 1–2%.</p><h3>Closing the Loop: Attribution</h3><p>The journey doesn’t end when the ad finishes. A plethora of ad tracker events are fired every second for qualifications, selections, impressions, and clicks from the moment the request lands on the server to the point when the ad is shown to the users on the app’s UI.</p><p>Our big-data pipelines process millions of these events every second, converting raw telemetry into accurate delivery metrics and advertiser dashboards. But delivery metrics alone aren’t enough. Advertisers ultimately care about outcomes.</p><p>This is where attribution comes in.</p><h4>Click-Through Attribution (CTA)</h4><p>Click-Through Attribution measures conversions that happen after a user <strong>clicks</strong> on an ad and later completes an action such as installing an app, signing up, or making a purchase.</p><p>User clicks -&gt; User converts -&gt; Conversion is credited to the ad.</p><h4>View-Through Attribution (VTA)</h4><p>This is like a subconscious win. Many users don’t click an ad during a tense IPL match (because FOMO!), but they might buy that product or install that app the next day. VTA tracks users who <em>saw</em> the ad and later converted, proving that the ad worked even without a click.</p><p>VTA captures the impact of brand exposure which helps when we need a large-scale awareness campaigns for some new product of a big brand.</p><h3>Conclusion</h3><p>So the next time you see a 30-second ad on JioHotstar, remember: you aren’t just watching a video. You are witnessing the finale of a sub 100-millisecond symphony!</p><p>It’s a world where algorithms balance the scales of supply and demand, where SGAI stitches a personalized experience for millions concurrently, and where big-data pipelines turn billions of pings into meaningful business insights.</p><p>For the customer, it’s just a break in the game, but for us, it’s one big complex dance.</p><p><em>Do you want to participate in building systems that handle intricate sub millisecond decisions? Do check out </em><a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering"><em>open roles</em></a><em> if you want to build for millions of customers, who will use features that you build!</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=e7b7f9921e46" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/journey-of-an-ad-request-the-hidden-engineering-e7b7f9921e46">Journey of an Ad Request: The Hidden Engineering</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Consistent Generative Video: Building Video Models That Remember]]></title>
            <link>https://blog.hotstar.com/getting-past-the-first-frame-building-video-models-that-remember-31f03d0710e1?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/31f03d0710e1</guid>
            <category><![CDATA[wan-2-1]]></category>
            <category><![CDATA[lora]]></category>
            <category><![CDATA[ltx]]></category>
            <category><![CDATA[generative-video]]></category>
            <dc:creator><![CDATA[Yaswanth Karnati]]></dc:creator>
            <pubDate>Wed, 01 Jul 2026 06:53:14 GMT</pubDate>
            <atom:updated>2026-07-01T07:01:44.225Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*2zWBxCjaDNWWL4npq5COgw.png" /></figure><h3><strong>Introduction</strong></h3><p>Generative AI offers a path to content creation at speeds and scales that traditional pipelines cannot match. One high-value use case: making existing IP characters-sports presenters, show hosts, brand ambassadors -deliver new scripts in their own likeness, without a reshoot, within IP &amp; contract limitations.</p><p>The premise is simple: take a reference image, provide audio, and generate a video where the character speaks naturally. In practice, this is harder than it sounds. Base proprietary video models like Kling, Veo often struggle with lip-sync quality, identity drift, and motion consistency.</p><p>To address this, we built a production pipeline for character-consistent talking-head generation using <strong>LoRA (Low-Rank Adaptation)</strong>, a lightweight fine-tuning method that adapts a base model to a specific character without retraining the entire network. This makes character adaptation practical to train, store, and deploy across multiple IP characters at scale.</p><h3><strong>Why LoRA?</strong></h3><p>LoRA works by freezing the original model weights and injecting small, trainable low-rank matrices into specific layers. Instead of updating a weight matrix W directly, LoRA decomposes the update into two smaller matrices: ΔW = A × B, where A and B have a rank r far smaller than the original dimensions. For rank 32 (our configuration), each adapted layer gains only ~0.1% additional parameters-producing a ~400MB adapter file vs. 41GB for the full model. Multiple character adapters can be swapped at inference without reloading the base model.</p><p>LoRA is well-established for image generation, but video introduces harder challenges: the adapter must preserve identity across dozens of frames (not just one), motion and structure are entangled across diffusion noise levels, and video LoRAs train on far fewer samples (15–20 clips vs. hundreds of images), making overfitting the dominant risk.</p><p>This blog covers how we fine-tuned two video architectures-Wan 2.1 and LTX-2-with LoRA, the data preparation that made it possible, and the results that determined which approach works at production scale.</p><h3>Model Training</h3><h4><strong>Preparing Training Data for Character LoRAs</strong></h4><p>The quality of a character LoRA depends on the quality of its training data, so our pipeline begins with a data preparation stage that converts raw footage into curated, LoRA-ready datasets. The system takes two inputs: source video footage and a clear reference image of the target IP character. As shown in <strong>Figure 1</strong>, the pipeline moves from scene preparation and character filtering to captioning, latent precomputation, and LoRA training.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/508/1*hEvslOtUdoZ_lwRqnubhKw.png" /><figcaption>Figure 1: Character LoRA training pipeline used in production</figcaption></figure><p>Raw videos are first split into scenes using <strong>PySceneDetect</strong>, with very short clips removed to avoid noisy transitions and montage fragments. We then use InsightFace / RetinaFace to detect faces and extract embeddings, which are matched against the reference image using cosine similarity. This retrieval stage achieved about 85% recall in identifying clips containing the target character. Only clips where the target character appears consistently are retained, even in multi-speaker or edited footage.</p><p>The retained clips are classified by shot type to preserve viewpoint balance and framing diversity. This helps the LoRA generalize better instead of overfitting to a single dominant angle.</p><p>The most important step is captioning, where each curated clip is described using <strong>Qwen2.5-Omni/Gemini Flash</strong> and assigned a unique trigger word for the target IP character. As illustrated in <strong>Figure 1</strong>, these captioned clips are then precomputed into reusable video latents and text embeddings before entering LoRA training.</p><p>At inference time, the trained adapter is invoked through the same character trigger, making the workflow reusable across prompts and scalable across multiple IP characters.</p><h4><strong>Model Selection</strong></h4><p>We focused on <strong>Wan 2.1</strong> and <strong>LTX-2</strong> because they were the strongest open-weight candidates in our evaluation, offering image-to-video quality closest to proprietary models according to <a href="https://artificialanalysis.ai/video/leaderboard/image-to-video?category=people">leaderboard ranking</a> while also supporting character-centric generation and LoRA-based adaptation in a practical training pipeline.</p><p>These two models also gave us useful architectural contrast. <strong>Wan 2.1</strong> was a strong initial candidate for character adaptation, while <strong>LTX-2</strong> offered a more flexible and production-oriented inference stack. Comparing them helped us assess quality, controllability, and operational fit for scalable deployment.</p><h4><strong>Experiment 1: Wan 2.1</strong></h4><p>Our first LoRA experiment was on Wan 2.1 I2V 14B. Training data was manually curated clips from out IP character catalogue.</p><p><strong>Single-Stage LoRA (Baseline):</strong> Single-stage LoRA trains one adapter across the full diffusion timestep range from 0 to 1000. This means the same adapter has to learn structure, motion, and fine detail together.</p><p>The limitation is that these aspects of generation are controlled by different timestep regions during denoising. As a result, the adapter learns a compromise rather than specializing effectively for any one of them.</p><p><strong>Observations: </strong><em>Over-articulated lips, robotic motion, visible jumps between clips, and identity drift.</em></p><p><strong>Three-Stage LoRA (Improved):</strong> Three-stage LoRA splits training into three adapters across different time-step ranges, allowing structure, motion, and fine detail to be learned separately, as shown in <strong>Table 1</strong>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/940/1*259lPBoaGVjeIS9T-brdLg.png" /><figcaption>Table 1: Three-stage LoRA training setup for Wan 2.1</figcaption></figure><p>Three separate LoRA files are loaded at inference with the low-noise adapter boosted to 1.3x scale for sharper lip definition. We also tuned inference: audio guidance 5.0 → 3.5 (curbs over-articulation), motion frames 9 → 15 (smoother transitions), text guidance 1.0 → 0.8 (better identity), sample steps 40 → 50 (higher quality).</p><p><strong>Observations: </strong><em>Mouth opening, clip transitions, and identity preservation all improved significantly. Though, lip movements remained unconvincing for production and motion consistency was weak. Wan 2.1 had reached its architectural ceiling.</em></p><h4><strong>Experiment 2: LTX-2</strong></h4><p>We then moved to LTX-2, a diffusion-based video generation model designed for efficient, high-quality image-to-video synthesis. Its architecture separates visual compression, text conditioning, and denoising into a cleaner training and inference stack, making it easier to adapt with LoRA for character-specific generation. In practice, this gave us a more flexible and production-friendly setup than Wan 2.1, with better support for stable identity preservation and higher-resolution outputs through two-stage upsampling from 960×544 to 1920×1152.</p><p>Unlike Wan 2.1, LTX-2 works well with a simpler single-stage LoRA setup. That reduces training complexity, simplifies checkpoint management, and makes it easier to operationalize across multiple IP characters. The pipeline keeps LoRA rank and optimization settings broadly aligned with Wan 2.1 for consistency, but the model itself requires less intervention to reach strong results.</p><h4><strong>Training Configuration (LTX-2):</strong></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/938/1*pOKe87AFEsWWqQItitmE0A.png" /><figcaption>Table 2: Training config. for LTX-2</figcaption></figure><p><strong>Overfitting Behavior: </strong>As shown in <strong>Figure 2,</strong> character features begin to emerge around 1k–3k steps, recognition becomes consistent by 3k–6k, and identity generalizes more reliably to unseen scenes in the 6k–9k range. Beyond 9k steps, the model begins to memorize training poses, causing outputs to become stiff and overly specific.</p><p>Based on this behavior, checkpoints were saved every <strong>3k steps</strong> and evaluated on unseen prompts, with the best production checkpoints typically selected from the <strong>6k–9k step</strong> range.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Vk57QL7tE1huP1zmDZHp0g.png" /><figcaption>Figure 2: LTX-2 Training Progression Across Checkpoints</figcaption></figure><p><strong>Observations: </strong><em>In general, the fine-tuned LTX-2 outputs showed stronger identity retention, more natural lip-sync, and smoother motion than the Wan 2.1 variants. It was also more reliable in carrying the character into unseen scenes, making it the strongest model in our pipeline from a production standpoint.</em></p><h3><strong>Results</strong></h3><p><strong>Model Comparison: Wan 2.1 vs. LTX-2 vs. Veo 3.1 vs. Kling 3.0</strong></p><p>To contextualize the LoRA results, we compared outputs from both fine-tuned models against <strong>Veo 3.1 and Kling 3.0</strong>, which was used without LoRA since it does not support adapter injection. All evaluations were conducted on IP characters from our catalog across both seen training scenes and unseen contexts. The quantitative and qualitative comparison across models is summarized in <strong>Table 3</strong>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*2laC2Hnb7LBMdF-HcvOQZg.png" /><figcaption>Table 3: Comparision of models across different evaluation metrics</figcaption></figure><p>As shown in <strong>Table 3</strong>, <strong>LTX-2</strong> provided the best balance of controllability, visual fidelity, and operational simplicity. It preserved identity more reliably across unseen contexts, produced more natural lip-sync, and fit more cleanly into a reusable character generation workflow. <strong>Wan 2.1</strong> remained a viable supported path, but it required more specialized tuning and delivered less consistent results. <strong>Veo 3.1 and Kling 3.0 </strong>demonstrated the strength of a large base model, but without LoRA support it was not a practical solution for reusable, character-specific generation across a catalog.</p><h3><strong>Conclusion</strong></h3><p>We built a LoRA fine-tuning pipeline that produces character-consistent talking head videos for JioHotstar IP characters at production quality. The core lessons: architecture matters more than tuning-Wan 2.1’s three-stage noise-level specialization improved results but could not overcome its architectural ceiling, while LTX-2 delivered larger gains with a simpler single-stage LoRA.</p><p>Data quality compounds-viewpoint symmetry and shot-type diversity in training data consistently improved generalization across characters. <em>Character selection shapes everything downstream-direct-to-camera monologue with clean audio produces the most reliable training and evaluation signal</em>.</p><p>In a follow-up post, we will cover our LLM-as-a-Judge evaluation framework and training data duration optimization experiments.</p><p><em>Do you want to contribute to the development of Indic language models? Do check out </em><a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering"><em>open roles</em></a><em> if you want to build for millions of customers, who will use features that you build!</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=31f03d0710e1" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/getting-past-the-first-frame-building-video-models-that-remember-31f03d0710e1">Consistent Generative Video: Building Video Models That Remember</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Beyond Imitation: Teaching AI to Speak Like a Character in Hindi]]></title>
            <link>https://blog.hotstar.com/beyond-imitation-teaching-ai-to-speak-like-a-character-in-hindi-476bd33f049a?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/476bd33f049a</guid>
            <category><![CDATA[chatterbox]]></category>
            <category><![CDATA[speech-models]]></category>
            <category><![CDATA[indic]]></category>
            <category><![CDATA[generative-ai-use-cases]]></category>
            <dc:creator><![CDATA[Sagar Tekwani]]></dc:creator>
            <pubDate>Fri, 08 May 2026 13:03:55 GMT</pubDate>
            <atom:updated>2026-05-08T13:03:54.112Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*A0dIMHLeWjzp4oCqjnjaVA.png" /></figure><h3>Introduction</h3><p>As JioHotstar serves a vast, diverse audience, generative AI is crucial for fast, scalable content creation across languages and formats, capabilities that traditional methods cannot match. Content systems are rapidly shifting to generative-first architectures, synthesising video, audio, and narrative programmatically.</p><p>However, progress is uneven: visual generation is advancing rapidly, while audio generation lags, especially for Indic languages, which lack the mature datasets and investment seen in English speech.</p><p>Under realistic production workloads in cinematic contexts, existing speech systems struggle to preserve emotional expressiveness, maintain pronunciation stability, and deliver consistent outputs across repeated generations.</p><p>These failures persist even on widely adopted commercial platforms, making audio the primary bottleneck in deploying AI-driven content pipelines alongside increasingly capable video systems.</p><h3>Why Indic Audio Breaks at Scale</h3><p>Speech generation systems that appear stable in demos often degrade under production-scale workloads. For Indic languages, these failures surface earlier and more acutely due to four structural constraints:</p><ul><li><strong>Lack of emotion-rich training data: </strong>Indic speech datasets are dominated by neutral narration. Unlike English, which benefits from large emotionally diverse corpora and mature conditioning techniques, Indic models struggle to produce emotionally aligned speech in character-driven settings.</li><li><strong>Insufficiency of prompt-based control: </strong>Prompting influences tone superficially but cannot enforce consistency or repeatability. Without explicit modeling of emotion, language-specific phonetics, and stability constraints, prompt-level approaches fail production reliability requirements.</li><li><strong>Linguistic complexity and code-switching: </strong>Hindi speech relies heavily on <a href="https://en.wikipedia.org/wiki/Prosody_(linguistics)">prosody</a>, vowel duration, and <a href="http://philseflsupport.com/weakening.htm">rhythmic stress</a>. Most systems lack language-aware mechanisms to resolve these variations, producing inconsistent or unnatural Hinglish output.</li></ul><h3>How was this Solved?</h3><p>We adopt a modular methodology that treats <strong>expressive speech as a core modelling challenge</strong>. Our approach combines carefully selected open-source architectures with targeted adaptation strategies for Hindi and Hinglish, augmented by language-aware preprocessing, emotion-centred data curation, and evaluation frameworks designed for production-scale content pipelines.</p><h4>Architecture Selection</h4><p>We evaluated three dominant architectural classes.</p><p>Non-autoregressive TTS models (e.g., IndicF5) offer low latency and stable pronunciation, but their explicit factorisation of prosody into auxiliary predictors limits expressiveness emotional variation collapses toward the training mean, producing perceptually flat speech.</p><p>Latent diffusion models (e.g., FishAudio) operate in compressed acoustic latent space and can capture global structure, but their effectiveness depends on pretraining coverage; lacking Indic pretraining, FishAudio showed high inference variability and was excluded from further benchmarking.</p><p>LLM-backed generative architectures emerged as the strongest candidate. These systems frame speech synthesis as conditional token generation, with text and reference audio jointly guiding the production of speech tokens.</p><p>Crucially, by learning prosody and rhythm <strong>implicitly through sequence modelling rather than through explicit auxiliary predictors</strong>, they are better positioned to capture character-specific expressive patterns making them the ideal foundation for character-consistent, emotionally aligned Hindi and Hinglish audio generation.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*7H-42r4icFaWqA11-0EhFA.png" /><figcaption><em>Fig.1: Arch. diagram for Chatterbox(Open-Source) Audio Model</em></figcaption></figure><p>As shown in Fig. 1, Chatterbox-class systems consume two primary inputs at inference time: the target text and a short reference audio clip. The reference audio is encoded by a <strong>Speech Tokenizer</strong> into discrete speech tokens and a speaker embedding, capturing voice timbre and coarse speaking characteristics.</p><p>These, combined with tokenized text, form the conditioning context for a large causal transformer (<strong>Text-to-Token, T3</strong>), which autoregressively predicts speech tokens. By conditioning jointly on text, speaker embeddings, reference speech tokens, and previously generated tokens, the T3 model captures long-range dependencies across linguistic content, pacing, and prosody allowing expressive patterns to emerge implicitly.</p><p>The predicted tokens are then decoded into continuous acoustic representations and converted to a waveform by a neural vocoder.</p><p>Key scientific advantages over alternatives:</p><p>(1) <strong>Implicit prosody modeling: </strong>rhythm, pacing, and emotional variation are learned through sequence modeling of speech tokens, not factored into explicit predictors that collapse under sparse expressive data.</p><p>(2) <strong>Strong conditioning pathways:</strong> text, style, and character embeddings are integrated early and jointly, preserving conditioning signals throughout generation and reducing drift across inference runs.</p><p>We conducted controlled experiments on Hindi scripts from our catalog evaluated model behavior across key dimensions. Metrics were computed over multiple inference runs to assess stability and repeatability rather than best-case performance. FishAudio was excluded due to the absence of Indic pre-training data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*g7_inY3QgMmJrp9P5N48QQ.png" /><figcaption><em>Table 1: Baseline comparison across architectural classes (Indic-F5 vs. Chatterbox)</em></figcaption></figure><h4><strong>Why Off-the-Shelf Inference Was Insufficient</strong></h4><p>Even with Chatterbox pre-trained on Hindi, zero-shot inference exposed a clear gap between voice timbre matching and expressive speaking style reproduction.</p><blockquote><strong>Generated speech consistently lacked character-specific prosody, rhythm, and emotional delivery producing perceptually monotonic outputs.</strong></blockquote><p>The root cause lies in the architecture’s conditioning design: reference audio provides coarse speaker and acoustic cues, while the T3 model learns an averaged prosody distribution across speakers during large-scale pre-training. This favours generalisation and stability but suppresses idiosyncratic speaking styles.</p><p>These effects are amplified for Hindi and Hinglish, where meaning and intent rely heavily on timing, stress, and intonation. In practice, zero-shot inference produced weak emotional alignment, pronunciation inconsistencies for mixed-language text, and high variability across runs establishing the need for targeted fine-tuning to explicitly encode expressive and language-specific priors.</p><h4>Data Preparation: Speech Enhancement and Emotion Annotation</h4><p>Noisy source audio is cleaned through a sequential pipeline: <strong>HTDemucs</strong> (speech isolation) → <strong>VoiceFixer</strong> (reverb/artifact removal) → <strong>DeepFilter</strong> (enhancement), or <strong>resemble-enhance</strong> as a single-pass alternative then gate-filtered via <strong>DNSMOS</strong> and loudness normalized.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qQXcdgDyczqj3h_MRI6oIA.png" /><figcaption><em>Fig-2: Speech Enhancement and Audio Processing Pipeline</em></figcaption></figure><p>For emotion labeling, audio segments are diarised, transcript-aligned, and classified by <strong>Qwen</strong> and <strong>Gemini</strong> against a predefined emotion taxonomy. Gemini was selected as the final annotator after outperforming Qwen on expressive classes (Happy/Energetic/Expressive) on a held-out validation set, directly translating to richer prosodic diversity in the training data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*IT2jH0yZImLsGoyrs1knqQ.png" /><figcaption><em>Table 2: Emotion annotation recall — Qwen vs. Gemini</em></figcaption></figure><p>Drawing on in-house archives and the enhancement pipeline, we curated a cross-lingual (Hindi + English) dataset spanning hundreds of speakers, deliberately designed to expose the model to code-switching and diverse prosodic patterns:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*gR4PKYvNPO--D0jtJQ_RGA.png" /><figcaption><em>Table 3: Fine-tuning dataset composition</em></figcaption></figure><h4>Fine-Tuning Strategy</h4><p>Fine-tuning was applied selectively to the <strong>Text-to-Token (T3)</strong> generative model, while keeping the speech tokenizer, decoder, and vocoder frozen. This design choice is scientifically deliberate: the T3 model needs to learn <strong><em>how</em></strong> the character speaks — pace, emphasis, and emotional cadence — without altering <strong><em>who</em></strong> the character sounds like.</p><p>However, since the voice identity is encoded in the frozen tokenizer and decoder, approximately one hour of expressive speech is sufficient to shift the T3’s conditional token distribution toward character-specific prosodic patterns, without triggering identity drift.</p><p>Training used the <strong>AdamW optimizer with a conservative learning rate</strong> to preserve pre-trained weights, a cosine learning rate schedule with a short warm-up for training stability, and <strong>gradient clipping</strong> to prevent unstable updates.</p><p>The training objective focused solely on predicting audio tokens from text inputs. <strong>Early stopping</strong> was implemented against validation loss to halt training before overfitting.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*W5B_o7dAxKAyVBs4OHDISg.png" /><figcaption><em>Fig. 3: Tensor-board training logs — training vs. validation loss</em></figcaption></figure><p>As shown in Fig.3, after 3–4 epochs, training loss continued declining toward zero while validation loss began rising, a clear overfitting signal, corroborated by a rise in gradient norm. <strong>The early stopping mechanism successfully arrested training before severe overfitting</strong>, saving the best-performing checkpoint and ensuring the final model generalizes well to unseen data.</p><h3>Results and Observations</h3><p>Fine-tuning outcomes were evaluated using quantitative metrics and systematic qualitative audits. Quantitative metrics captured pronunciation accuracy, emotional alignment, long-form consistency, and speaker similarity across multiple inference runs. Qualitative audits identified failure modes not fully reflected by aggregate metrics, particularly for Hinglish and domain-specific content.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*zXSoJixRJzcGmtD1vRk52A.png" /><figcaption><em>Table 4: Fine-tuned model performance vs. base model and 11labs benchmark</em></figcaption></figure><p>The most scientifically significant result is the <strong>+27% gain in Speaker Similarity</strong> despite the frozen decoder and vocoder. This confirms that T3 fine-tuning successfully learned character-specific prosodic priors <strong>without destabilizing voice identity</strong> — validating the selective fine-tuning hypothesis.</p><p>Emotion Alignment improved by 19% and Pronunciation Accuracy by 23%, with Mean Opinion Score (MOS) rising from 3.3 to 4.2 (approaching the 11labs benchmark of 4.3). The following issues were observed during our extensive testing, which were mitigated as described:</p><h4>Issue-1: Gibberish Artifacts in Short Phrase Generation</h4><p><strong>Observation: </strong>For utterances under five seconds, the fine-tuned model occasionally generated non-linguistic or gibberish speech toward the end of the audio. This was reproducible for a subset of short phrases and was not observed in the pre-trained base model.</p><p><strong>Interpretation: </strong>Short sequences provide fewer autoregressive steps, increasing sensitivity to distribution shifts introduced during fine-tuning. The model overshoots termination boundaries or produces unstable token transitions near sequence end, a failure mode specific to the fine-tuned distribution.</p><p><strong>Mitigation: </strong>Root cause analysis identified the failure at the generation boundary level. Targeted fixes were applied during data curation (ensuring short-phrase samples had clean terminal boundaries) and at the generation stage (constraining end-of-sequence token behavior). After applying these mitigations, gibberish artefacts were no longer observed in short-phrase inference.</p><h4>Issue-2: Pronunciation Errors in Domain-Specific and Hinglish Vocabulary</h4><p><strong>Observation: </strong>Despite overall natural speech quality, domain-specific terms and code-switched vocabulary were mispronounced even when the base model had been pre-trained on Hindi and English appeared sporadically in training audio.</p><p><strong>Interpretation: </strong>The root cause is ambiguity in text-to-phoneme alignment under code-switching. The model had largely seen fully Hindi or fully English contexts during training, whereas inference-time inputs (e.g., cricket commentary) contained tightly interleaved Hindi and English tokens, leading to unstable pronunciation for certain words.</p><p><strong>Mitigation: </strong>A pre-inference transcript normalisation strategy using a lightweight lexicon mapping replaced consistently mispronounced words with acoustically preferred alternatives before inference:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*T7ZdAQDRlawFm2-8MS4SMA.png" /><figcaption><em>Table 5: Sample pronunciation normalization mappings</em></figcaption></figure><p>This normalization produced clear improvements in pronunciation consistency without affecting overall fluency.</p><h3>Conclusion</h3><p>By combining LLM-assisted emotion annotation with targeted T3 fine-tuning on enhanced expressive data, we achieved <strong>~20% improvement in Emotion Alignment, ~27% in Speaker Similarity, and MOS rising from 3.3 to 4.2, </strong>approaching commercial benchmarks while maintaining stability across repeated inference runs.</p><p>This establishes a reliable foundation for character-driven Hindi and Hinglish audio generation at scale. In the next blog, we extend this work to regional Indic languages, tackling even steeper challenges in data scarcity and linguistic diversity.</p><p><em>Do you want to contribute to the development of Indic language models? Do check out </em><a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering"><em>open roles</em></a><em> if you want to build for millions of customers, who will use features that you build!</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=476bd33f049a" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/beyond-imitation-teaching-ai-to-speak-like-a-character-in-hindi-476bd33f049a">Beyond Imitation: Teaching AI to Speak Like a Character in Hindi</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Beyond Manual Checks: How We Built a 24/7 “Watchdog” for Live Sports Streaming]]></title>
            <link>https://blog.hotstar.com/beyond-manual-checks-how-we-built-a-24-7-watchdog-for-live-sports-streaming-026fdcaebc5e?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/026fdcaebc5e</guid>
            <category><![CDATA[video-validation]]></category>
            <category><![CDATA[automation]]></category>
            <category><![CDATA[live-streaming]]></category>
            <dc:creator><![CDATA[Abhishek Sinha]]></dc:creator>
            <pubDate>Wed, 08 Apr 2026 12:43:23 GMT</pubDate>
            <atom:updated>2026-04-08T12:43:21.779Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*McegzCiz7HRLm4EkvndLqQ.png" /></figure><p>In the world of live sports streaming, there is no “undo” button. Live events bring out emotions in viewers and any minor buffering or failure can cause disappointment in our customers, especially if it happens in crunch moments of the live event.</p><p>Our Live video validation has gradually moved from being fully manual to islands of automation. We had the tools to check manifests and APIs, but we lacked a cohesive brain to run them. Here we talk about how we moved from manual intervention to an automated, cohesive, real-device validation ecosystem.</p><h3>Manual “QA” Failed at Scale</h3><p>Before we built our current pipeline, our validation process was a series of stretch efforts that simply couldn’t keep up with the scale of modern streaming. Manual QA can work for a single video and small number of matches, but when we look at streaming a live event in 12 languages, 5 camera angles, 74 matches over 2 months, automation is the only way to reliably run a QA process. At peak, JioHotstar (JHS) has seen close to 55 concurrent live streams, each with their unique stream variations.</p><ul><li><strong>The Orchestration Gap:</strong> We started automated stream checks early on, but stream checks were started manually. If an engineer forgot to trigger a script for a specific regional feed, that feed went unmonitored.</li><li><strong>The “Orphaned” Test Problem:</strong> We would often find validation scripts running for hours after a match ended because there was no central logic to tell them to stop. Automation is wonderful but cloud costs are real too.</li><li><strong>The Silent Failure:</strong> A manifest might look perfect on a server (<strong>200 OK</strong>), but the actual player on a Smart TV might be stuck on a buffering wheel due a media sequence mismatch or due to ads related discontinuity tags or a DRM mismatch.</li></ul><p>We didn’t just need a bigger QA team; we needed an <strong>orchestrator for our automation, a deeper video validation mechanism and real devices</strong>.</p><h3>Phase 1: The Steering (Video Validation Orchestrator)</h3><p>It’s not enough for a manual or automated test to certify that a live event is working by testing for a couple of seconds. The automation needs to follow the live event’s lifecycle. The same video stream needs to be continuously validated throughout the event.</p><p>Apart from languages and camera angles we discussed earlier, JioHotstar’s video encoding systems also generate many video variants based on device (e.g. TV and mobile get different videos), CDN (what if a video is working on one part and failed on another), advanced video and audio experience like Dolby Vision, Dolby 5.1, HDR10, 4K etc. These video streams can fail independently and hence must be tested independently too.</p><p>Instead of manual triggers, the Orchestrator treats our <strong>Content Metadata System (CMS) </strong>as the absolute source of truth and also validates all the video variants.</p><h3>How the “Steering” Works</h3><p>The Orchestrator operates on a <strong>Discovery-and-Dispatch</strong> model:</p><ul><li><strong>State Discovery:</strong> It polls the metadata provider to maintain a real-time map of every scheduled match, regional feed, and alternative camera angle (like VR360). We are already working on consuming content publish/un publish events to make this process even better.</li><li><strong>Automated Lifecycle Management:</strong> It automatically “spins up” validation workers before “match-start” and, more importantly, “spins them down” once the content is unpublished to prevent resource leakage.</li><li><strong>Parallel Execution:</strong> For every live event, the Orchestrator dispatches two distinct, concurrent validation strategies; server-side (Video Validator) and real-device (Play Portal)</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*SMVyjyjP4_O1M5lR9wpTGw.png" /><figcaption>video-streaming-validation-architecture</figcaption></figure><h3>Phase 2: Dual-Layer Validation</h3><p>To ensure total coverage, we split our focus between the “plumbing” of the stream and the actual “eyes” of the user.</p><h3>Strategy A: Video Validator (Deep-Stream Server Side Inspection)</h3><p>This is our <strong>Server-Side</strong> watchdog. It doesn’t look at the UI; it looks at the raw data packets and manifests to identify backend failures before they reach the player.</p><ul><li><strong>Manifest Integrity:</strong> Checks for discontinuities, missing segments, and media sequence drift.</li><li><strong>Ad-Tech Health:</strong> Validates <strong>SSAI</strong> (Server-Side Ad Insertion) markers and ensures <strong>CUE-IN/CUE-OUT</strong> timings are structurally sound.</li><li><strong>CDN Load-Balancing:</strong> Verifies that all CDN edges (Akamai, Cloud front, JIO etc.) are serving identical, healthy chunks</li></ul><p>Video Validator service is the first one to catch any issues with CDNs, encoders, and SSAI manifest stitchers. Any errors caught here are shown on the dashboard (covered later) and relevant alerts are raised for our teams.</p><p>Our players and CDNs are tuned to retry and recover from intermittent failures (like the ones in the diagram below), but we still need to know and understand the root cause when these failures happen.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*OD-DreU9cAyvhPcYPpQBWg.png" /></figure><h3>Strategy B: Play Portal (Real-Device Synthetic User Simulation)</h3><p>While the Validator checks the plumbing, <strong>Play Portal</strong> acts as the eyes. It launches automated playback sessions on <strong>real devices</strong> (Android Mobile, iPhones, TVs, Web Browsers) to ensure content is actually playable.</p><ul><li><strong>Playback Handshake:</strong> Does the player initialize and reach “playing” state within the SLA?</li><li><strong>UI Overlay Checks:</strong> All expected audio languages are listed in the player and is playable</li><li><strong>Buffering: </strong>Any rebuffering or other esoteric scenarios where video frames get stuck are detected here.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*r6YlfJVo-egaj1bhMdLe0Q.png" /></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ANuRL0hmYghxjUhwM5xkxA.png" /></figure><h3>The Pulse of the Stadium on a Dashboard</h3><p>When we monitor at this scale, the Grafana dashboards stop looking like lines of code and start looking like a heartbeat. We can actually track the stadium’s atmosphere through our validation logs:</p><ol><li><strong>The “Toss” Mini-Spike:</strong> 30 minutes before the first ball, our Orchestrator triggers a surge of validation workers. The dashboard “wakes up” as we see the first successful playback handshakes across 12 regional languages.</li><li><strong>The First Ball Explosion:</strong> As the first ball is bowled, request volume explodes. Our <strong>Kubernetes HPA</strong> (Horizontal Pod Autoscaler) kicks in, spinning up more Video Validator pods to handle the massive concurrency.</li><li><strong>The Middle-Overs Lull:</strong> During the middle overs, we look for long-running stability — monitoring for “Silent Drifts” or memory leaks in our real-device probes that only appear after 60 minutes of playback.</li><li><strong>The Second Innings “Re-Sync”:</strong> As the chase begins and a new wave of viewers joins, we ensure mid-roll ad markers are firing correctly, catching “Empty VAST” errors before they impact revenue.</li><li><strong>The Pre-Death Overs Ramp:</strong> In the final 5 overs, we are often using the last ounce of CDN bandwidth available in the country. Validation here hits the same intensity that the match is experiencing.</li><li><strong>The Show Continues(Match End):</strong> The match may end but our viewers are still here to find out who got the player of the match trophy, our validators continue to hum after the match.</li><li><strong>The Sudden Drop (Unpublish),</strong> Traffic drops, the Orchestrator detects the “Unpublished” state, and begins the graceful scale-down of our workers.</li></ol><p><strong>Engineering Takeaway:</strong> Watching this cycle isn’t just about “uptime.” It’s about <strong>Elasticity.</strong> Our system needs to be as dynamic as the game itself — expanding its “lungs” during the death overs and breathing out when the stadium goes dark.</p><h3>Real-World Impact: Detection in Action</h3><p>Our automated validations start much before toss and this has saved us from major regional outages by catching issues before we published matches for the audience:</p><ul><li><strong>Manifest 404s:</strong> Caught ingest gaps where regional feeds hadn’t reached the CDN edge, allowing for a fix before the first viewer clicked play. We spin up 100s of encoders for every match when one fails, it’s detected using multiple mechanisms including Video Validator.</li><li><strong>SSAI / SGAI Failures:</strong> Identified invalid cues during high-traffic windows, preventing “black screen” ad breaks.</li><li><strong>Media Sequence Drift:</strong> Detected sync issues across different bitrates that would have caused infinite buffering for users. This detection led to our player teams strengthening media sequence handling.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*e80pgemVOwlXPWQqKMNB8Q.png" /></figure><h3>The Takeaway: Building Confidence, Not Just Code</h3><p>Moving to an automated stream validation model wasn’t just a technical upgrade; it was the only way to operate to ensure that we caught problems before hearing about it from social media<strong>.</strong></p><p>By combining Server-Side Validation with Client-Side Automation, we created a safety net that covers the entire <strong>“journey of a pixel.”</strong></p><p>We no longer hope the stream is working; we know it is, from the moment it leaves the stadium to the moment it renders in the fan’s living room.</p><p><em>Are you interested in ensuring you build video quality agents? Do check out </em><a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering"><em>open roles</em></a><em> if you want to build for millions of customers, who will use features that you build!</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=026fdcaebc5e" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/beyond-manual-checks-how-we-built-a-24-7-watchdog-for-live-sports-streaming-026fdcaebc5e">Beyond Manual Checks: How We Built a 24/7 “Watchdog” for Live Sports Streaming</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[AI Video Generation at Scale — Prompts Don’t Scale]]></title>
            <link>https://blog.hotstar.com/ai-video-generation-at-scale-prompts-dont-scale-640c93778df9?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/640c93778df9</guid>
            <category><![CDATA[ai-video-prompting]]></category>
            <category><![CDATA[ai-evals]]></category>
            <category><![CDATA[generative-video]]></category>
            <dc:creator><![CDATA[Kapil Gupta]]></dc:creator>
            <pubDate>Sat, 28 Mar 2026 15:33:03 GMT</pubDate>
            <atom:updated>2026-03-28T15:38:18.687Z</atom:updated>
            <content:encoded><![CDATA[<h3>AI Video Generation at Scale — Prompts Don’t Scale</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*jMQ2kjrcio-pHheVYYnGcA.png" /></figure><p>AI video generation has rapidly moved from a mere novelty to a practical tool, allowing users to easily create brief video clips from text prompts. While generating a single, impressive video is straightforward, the true difficulty lies in scaling this process to produce consistent, repeatable, and production-grade content.</p><p>This post contends that the inability to generate high-quality, production-grade AI content at scale is due to systemic failures in the current approach, not limitations in the core AI models themselves.</p><h3>The Requirements of Production Quality</h3><p>Achieving production-grade video demands a system capable of handling several stringent requirements:</p><ul><li><strong>Volume:</strong> Generating a high quantity of short videos, not just isolated clips</li><li><strong>Complexity:</strong> Supporting multiple distinct scenes within a single video</li><li><strong>Consistency:</strong> Reliably reusing and maintaining characters, locations, and objects across all content</li><li><strong>Collaboration:</strong> Facilitating team workflows rather than being restricted to individual use</li><li><strong>Predictability:</strong> Ensuring stable cost structures and reliable delivery timelines</li><li><strong>Compliance:</strong> Adhering to brand guidelines, safety regulations, and intellectual property rights</li></ul><p>It is against these rigorous standards that the inherent flaws of existing AI video systems become obvious. The core challenges preventing production-scale AI content are detailed below:</p><h3>#1: Prompt-Centric Systems Cannot Scale</h3><p>While prompting is an excellent interface for exploration, it proves fragile for production work. The prompt-based approach suffers because:</p><ul><li>Consistency is solely dependent on the user’s conceptualisation and is not encoded within the system</li><li>Output quality fluctuates based on the specific user who wrote the prompt</li><li>Re-running an identical prompt offers no guarantee of the same outcome</li><li>Operational knowledge remains uncodified, failing to be formally integrated into the system</li></ul><p>As teams and projects expand, prompts inevitably become:</p><ul><li>Unnecessarily long</li><li>Managed through error-prone copy-and-paste practices</li><li>Subject to small, undocumented alterations</li><li>Difficult to systematically analyze and maintain</li></ul><p>At scale, this leads to <strong>prompt drift</strong>, where video outputs increasingly diverge from the original creative intent, despite the seemingly consistent prompts. Prompting is fundamentally an improvisational tool, not a structured system contract.</p><h3>#2: Continuity Failures: Drifting Characters, Mutating Scenes, Disappearing Objects</h3><p>In professional video, continuity is critical: a character’s appearance and voice must be uniform, locations must not subtly change without authorization, and props must persist as intended.</p><p>Yet, in most AI video generation:</p><ul><li>Characters are fully regenerated for every instance</li><li>Locations are merely implied, lacking systematic enforcement mechanisms</li><li>Objects are treated as incidental details instead of continuously tracked entities</li></ul><p>This results in a gradual breakdown of realism, manifesting as:</p><ul><li>Shifts in facial structure</li><li>Subtle changes in wardrobe</li><li>Erratic transformation of background environments</li><li>Minor but noticeable deviations in voice from the baseline</li></ul><p>Human viewers instantly perceive these inconsistencies, even if computational models score the outputs favorably. Consistency is not a minor visual detail; it is a fundamental cognitive contract with the audience.</p><h3>#3: Quality Control Is Manual and Subjective</h3><p>The quality assurance for most AI video generation relies heavily on:</p><ul><li>Human visual inspection</li><li>Subjective assessment (“This looks acceptable”)</li><li>Creative intuition, lacking objective metrics</li></ul><p>This method works only at low volumes. At production scale, it immediately introduces problems such as:</p><ul><li>Inconsistent judgment across human reviewers</li><li>Unpredictable drift in quality standards</li><li>The inadvertent release of suboptimal content</li><li>The incorrect rejection of high-quality outputs</li></ul><p>If output quality cannot be measured objectively, it cannot be reliably enforced. Production systems must integrate objective evaluation as a primary, foundational capability, not a secondary add-on.</p><h3>#4: Regenerations Lead to Unpredictable Cost</h3><p>When an error or undesirable outcome occurs in current AI video generation, the standard solution is simply:</p><p><em>“Try generating it again.”</em></p><p>AI video generation is inherently resource-intensive, particularly in terms of GPU utilization. While cost is a minor concern at low volumes, it becomes a critical issue at production scale. The common cost challenges include:</p><ul><li><strong>“Retry storms”:</strong> Cycles of repeated, unsuccessful retries</li><li>Unconstrained and potentially endless regeneration attempts</li><li>Over-reliance on expensive, high-specification models</li><li>A lack of correlation between production expenditure and final output quality</li></ul><h3>A Necessary Shift in Mental Model</h3><p>To overcome these systemic failures, effective production systems must prioritize:</p><ul><li>The system architecture first</li><li>Consistency and continuity first</li><li>Evaluation and quality control first</li></ul><p>To operate effectively at scale, AI video systems require a fundamental conceptual shift:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QTACmaz3P-fbknmyJvrccA.png" /></figure><p>This transformation necessitates a rigorous architectural approach, not merely computational workarounds.</p><p><em>Subsequent articles in this series will explore the crucial paradigm shifts required to establish AI video as a viable asset for professional teams.</em></p><p><em>Are you interested in AI Video Generation? Do check out </em><a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering"><em>open roles</em></a><em> if you want to build for millions of customers, who will use features that you build!</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=640c93778df9" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/ai-video-generation-at-scale-prompts-dont-scale-640c93778df9">AI Video Generation at Scale — Prompts Don’t Scale</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Orchestrating JioHotstar Traffic: The Difference Between a Loading Spinner and a Winning Six]]></title>
            <link>https://blog.hotstar.com/orchestrating-jiohotstar-traffic-the-difference-between-a-loading-spinner-and-a-winning-six-8a385b01380e?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/8a385b01380e</guid>
            <category><![CDATA[cdn]]></category>
            <category><![CDATA[technology]]></category>
            <category><![CDATA[live-streaming]]></category>
            <category><![CDATA[qos]]></category>
            <category><![CDATA[last-mile-delivery]]></category>
            <dc:creator><![CDATA[Karan Kaul]]></dc:creator>
            <pubDate>Mon, 09 Mar 2026 09:41:53 GMT</pubDate>
            <atom:updated>2026-03-09T09:41:52.459Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/767/1*IxovPCvikcIuEEHqKwz8FQ.jpeg" /></figure><p>If you’ve ever streamed a high-stakes match on JioHotstar, the experience probably felt like a simple “tap and play”. It’s a seamless transition from the app icon to the stadium; a flick of a finger and you’re right there, cheering for your players to hit another six. Beneath that simple play button lies one of the most aggressive engineering challenges in the world.</p><p>At Hotstar, we navigate a “<a href="https://blog.hotstar.com/t-for-tsunami-dealing-with-traffic-spikes-c22443bcdd3e"><strong>Tsunami</strong></a>” of traffic across one of the most complex network landscapes in the world. During events like the IPL or a high-stakes ODI match, we don’t just manage millions of users - we manage millions of unique network realities. If you thought just placing a Content Delivery Network (CDN) makes the magic happen — read on.</p><h3>The Illusion of the Monolithic Network</h3><p>Traditionally, CDN traffic management happens at the <strong>network provider </strong>or <strong>state</strong> level. For years, this was enough. But at our scale, “good enough” is the enemy of the audacious. A live event requires upwards of 60–80 Tbps of network bandwidth for streaming. To put things into perspective, it’s like instantly grabbing the <strong>entire 4K movie collection</strong> from JioHotstar (hundreds of blockbusters). Now imagine finishing all those downloads in just one second ….and then doing it again the<strong> next second and the next</strong>. That is the relentless pace of the tidal wave we face during a major event.</p><p>However, <strong>India’s network isn’t a monolith</strong>. It is a <strong>mosaic</strong> of fiber, 4G, 5G and fluctuating bandwidth that changes from one street to the next. A user on a 5G connection in South Delhi faces a vastly different network topology than a user on a local ISP in rural Rajasthan.</p><p>When you are only as good as the video you deliver, you realize that macro-level routing hits a ceiling. To provide the best <a href="https://en.wikipedia.org/wiki/Quality_of_service">Quality of Service</a> (QoS), we needed to treat every state, every city, and eventually every cohort, as a unique routing decision.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*NWpNsHiYvtc4oH3vqFuIBw.png" /></figure><h3>The Solution: The QoS Routing Manager</h3><p>To solve the “Last Mile” problem, we built a real-time observability and orchestration engine: the <strong>QoS Routing Manager</strong>.</p><p>The mission of this service is simple but massive: Observe crucial video metrics at a granular level and adjust traffic weights dynamically to ensure every user is mapped to the best possible CDN for their specific location.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Ia3iObTclEtPF3kL5rwbWQ.png" /></figure><h4>The Granular Cohort</h4><p>Instead of routing by “Maharashtra” or “Jio,” we segment users into <strong>Cohorts</strong>. A cohort is a group of people who share a common characteristic over a given period. Currently for us cohort is a specific combination of geographical, network and business categorisation i.e<strong> </strong><a href="https://www.cloudflare.com/learning/network-layer/what-is-an-autonomous-system/"><strong>ASN</strong></a><strong>-Country-State-City-UserType</strong>.</p><p>This allows us to detect if a specific provider is facing issues in a specific city, even if their national health looks perfect. By slicing the data this way, we can bypass localized congestion before it affects the broader user base</p><h4>The Power of The Scoring Logic</h4><p>The QoS Routing Manager pulls real-time performance data from our sophisticated in-house <a href="https://en.wikipedia.org/wiki/Online_analytical_processing">OLAP</a> analytical processing beast - called <strong>ARGUS, </strong>which serves as our eyes and ears when it comes to video performance across the platform. It collects heartbeat data from the devices, processes them and provides the telemetry to our service.</p><p>We map every metric - <strong>Playback Failure Rate (PFR)</strong>, <strong>Rebuffering</strong>, and <a href="https://www.cloudflare.com/learning/cdn/glossary/round-trip-time-rtt/"><strong>RTT</strong></a><strong> Latency </strong>- into normalized scores using a <a href="https://en.wikipedia.org/wiki/Piecewise_linear_function"><strong>Linear Piecewise Scoring</strong> function</a>. This allows us to define “Severity Buckets” (Ideal, Baseline, Sev3, Sev2, Sev1) based on direct business impact.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*h0MVZwlj9tjN5kSZelKx8w.png" /></figure><p>To determine which CDN “wins” for a specific cohort, we calculate a <strong>Cumulative Health Score</strong> based on a specific precedence order:</p><blockquote>Cumulative Score= ( X * PFR{score} ) + ( Y * Rebuffer{score} ) + ( Z * RTT{score} )</blockquote><p>We weigh PFR most heavily to ensure that “<strong>reachability</strong>” is the absolute priority, followed closely by the “<strong>fluidity</strong>” of the stream (Rebuffering) and the “<strong>snappiness</strong>” of the connection (RTT).</p><h4>Filtering the Noise: The Power of EWMA</h4><p>Raw network telemetry is inherently noisy. A momentary 4G tower hand-off in Jammu or transient packet loss in Bangalore can look like a critical failure in a 10-second window.</p><p>If our routing engine reacted to every micro-spike, we would introduce dangerous volatility, creating a “jittery” experience where users are constantly bounced between CDNs. We needed a way to smooth out the true performance trend from the momentary noise.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/699/1*XBPEBQmQXYJ-2bwLdKfiDg.png" /><figcaption>CDN Scores for different cohorts — observe the intermittent drops and spikes</figcaption></figure><p>To achieve this, we don’t pass raw metrics directly into our scoring engine. Instead, we pass all incoming telemetry through an <a href="https://corporatefinanceinstitute.com/resources/career-map/sell-side/capital-markets/exponentially-weighted-moving-average-ewma/"><strong>Exponentially Weighted Moving Average</strong></a><strong> (EWMA)</strong> filter.</p><p>The formula we use is:</p><blockquote>EWMA{t} = α * x_t + ( 1 — α ) * EWMA{t-1}</blockquote><p><em>where </em>EWMA<em>{t} is the new smoothed value, x_t is the raw input score, and </em>EWMA<em>{t-1} is the previous smoothed history.</em></p><p>We tune our smoothing factor <strong>α</strong> to approximately <strong>0.6</strong>. In practical terms, this means our engine ensures recent metrics have a stronger influence on the final score while keeping the <strong>last 5 scores to have significance</strong></p><p>Only once the metrics are smoothed by EWMA do we move to the routing.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Rpe7WzuMZdVqI4rA8XoPbA.png" /><figcaption>EWMA CDN Scores for same cohorts on network — smoother spikes</figcaption></figure><h4>Two-Phase Capacity Steering</h4><p>A high score isn’t the only requirement for routing. We must balance “Customer Joy” with “Infrastructure Dynamics”. For a media streaming entity, the biggest tradeoff is quality and available network bandwidth.</p><blockquote>It’s like a bridge where every viewer wants to drive a wide luxury bus (<strong>High Quality</strong>), but the physical lanes (<strong>Bandwidth</strong>) are finite; during a peak surge, there simply isn’t enough pavement to let everyone drive a bus at once without the bridge failing, so you have to balance the vehicle size just to keep everyone moving.</blockquote><p>Our engine keeps track of the used bandwidth on the CDNs and employs two distinct strategies:</p><ul><li><strong>Pre-Threshold Guardrail (&lt; X% Utilization):</strong> We preemptively throttle CDNs that are on track to hit their capacity limits too early, even if they are performing well.</li><li><strong>Uniform Exhaustion (&gt;= X% Utilization):</strong> During a surge, we shift logic to ensure all CDNs exhaust their capacity at the same time, squeezing every possible megabit out of our infrastructure.</li></ul><h4>Safety Gates: Resilience Over Risk</h4><p>In a system of this scale, “no update” is better than a “bad update.” We built a defense-in-depth approach covering both data integrity and routing logic.</p><ol><li><strong>ASN Constraints (Hard Binding):</strong> Physics and contracts still matter. Some CDNs only have presence on specific networks. Before any scoring happens, the system applies hard constraints to ensure we never route a cohort to a CDN that physically cannot serve it.</li><li><strong>Volatility Dampening (Max Deviation):</strong> To prevent wild swings in traffic that could de-stabilize the network, we cap the maximum percentage change allowed in a single iteration (e.g., a CDN cannot gain or lose more than y% share in one minute).</li><li><strong>The “Warm-Up” Floor (Minimum Weight):</strong> We never let a functional CDN drop to 0% traffic. We enforce a configurable floor (typically 5%). This keeps CDN caches warm and DNS paths active, ensuring that if we need to fail-back to them instantly during a crisis, they are ready to take the load immediately.</li></ol><h3>The Audacious Impact</h3><p>The results of moving to granular, cohort-based management has been significant. During a recent <strong>T20 Match</strong>, we ran a A/B rollout of the QoS Routing Manager with only ASN-State cohorts. For the treatment group:</p><ul><li><strong>Playback Failure Rate (PFR)</strong> improved by a staggering <strong>11%</strong>.</li><li><strong>Rebuffering</strong> and <strong>RTT Latency</strong> both saw a <strong>2%</strong> improvement.</li></ul><p>In our world of 50M+ concurrent users, an 11% improvement in PFR represents millions of users who stayed connected to the game instead of seeing a loading spinner. This is how we ensure that whether you are in a high-rise in Chennai or a village in Sikkim, you enjoy each and every six with best in class quality.</p><p>The work continues to build on the millions of datapoints that stream in that allow us to steer all our customer sessions to a stable viewing experience!</p><p><em>Are you interested in solving high-concurrency challenges at the edge? Do check out </em><a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering"><em>open roles</em></a><em> if you want to build for millions of customers, who will use features that you build!</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=8a385b01380e" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/orchestrating-jiohotstar-traffic-the-difference-between-a-loading-spinner-and-a-winning-six-8a385b01380e">Orchestrating JioHotstar Traffic: The Difference Between a Loading Spinner and a Winning Six</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Monetisation : Multi-period DASH x ExoPlayer]]></title>
            <link>https://blog.hotstar.com/monetisation-multi-period-dash-x-exoplayer-db8e8c00e521?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/db8e8c00e521</guid>
            <category><![CDATA[dash]]></category>
            <category><![CDATA[exoplayer]]></category>
            <category><![CDATA[streaming]]></category>
            <category><![CDATA[ads]]></category>
            <dc:creator><![CDATA[Abhishek Bansal]]></dc:creator>
            <pubDate>Wed, 04 Mar 2026 04:18:06 GMT</pubDate>
            <atom:updated>2026-03-04T04:18:04.447Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6OlEGvI87XY7LSE_hoNFiQ.png" /></figure><blockquote>Client Side Ad Insertion(CSAI) levels up using multi-period DASH in the Android ecosystem. This is our journey to upgrade ExoPlayer to leverage multi-period DASH.</blockquote><h3>Introduction</h3><p><a href="https://ottverse.com/single-period-vs-multi-period-dash/">Multi-period DASH</a> is a variant of <a href="https://en.wikipedia.org/wiki/Dynamic_Adaptive_Streaming_over_HTTP">DASH</a> format that offers significant advantages, such as the ability to insert dynamic content like disclaimers, dub cards, and ad breaks without re-encoding the entire video. This flexibility is crucial for delivering personalized and localized content to millions of users across diverse regions.</p><p>While the DASH standard natively supports multi-period manifests, the Android <strong>ExoPlayer</strong> ecosystem (specifically the AdsMediaSource component) was designed with a single-period assumption. This missing piece in the puzzle meant we couldn&#39;t support Client-Side Ad Insertion (CSAI) on these modern streams out of the box.</p><p>This post details our engineering journey: identifying the constraints, evaluating architectural alternatives, and ultimately redesigning ExoPlayer’s Ad handling to support multi-period content seamlessly.</p><h3>Context</h3><h4>The Evolution of DASH Content</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MxXNu-p6Y4_x3VQqXe9Nlg.png" /></figure><p>JioHotstar relied on single-period DASH for entertainment content, where the entire stream including ads was encapsulated within a single period. This approach limited flexibility: inserting disclaimers, dub cards, and to some extent, dynamic ad breaks required re-encoding the entire video, which was both time-consuming and resource-intensive.</p><p>Consider a common localization scenario: the same title is launched across multiple regions with different primary languages. Audio dubs may be recorded by different artists per market, so dub-card credits vary by region; likewise, legal and regulatory requirements differ, driving region-specific disclaimers.</p><p>Under a single-period workflow, each of these variations would necessitate a distinct transcode of the full asset. If we target ~20 regions, that implies ~20 full transcodes, leading to long time-to-market, high cost, and significant operational complexity — super unscalable.</p><p>Multi-period DASH addresses these challenges by dividing content into multiple periods, each representing a distinct segment (e.g., disclaimer, main content, ad break). This modularity enables seamless insertion of additional content without re-encoding, significantly reducing operational overhead. In practice, we can attach region-specific disclaimers as a lightweight period and swap dub-card credits per locale without touching the main essence, keeping a single <a href="https://medium.com/freelance-filmmaker/intermediate-codec-using-mezzanine-video-formats-a28a53d3256e">mezzanine</a> and avoiding redundant full transcodes.</p><h4>The Challenge of Ad Monetization</h4><p>While multi-period DASH brought flexibility, it also introduced a critical challenge to our Ad insertion mechanism for Video on Demand(VOD) content. ExoPlayer, our Android media player, lacked native support for client-side ad insertion in multi-period DASH streams.</p><h4>Root Cause</h4><p>ExoPlayer’s code restricted ad playback to single-period content, for Ads supported users, Exoplayer uses AdsMediaSource for Ads playback and tracking. The AdsMediaSource does not allow Multi period contents with Ads. It throws IllegalArgumentException if the content has more than one period.</p><p>The crash happens as soon as AdsMediaSource object is created, irrespective of whether there are actual ads inserted in the stream or not.</p><p>Digging into the library code, we found explicit assertions guarding against complexity:</p><pre>// Logic inside AdsMediaSource / SinglePeriodAdTimeline <br>Assertions.checkState(periodCount == 1);</pre><p>An immediate thought here would be to upgrade to latest <a href="https://github.com/androidx/media">AndroidX Media3</a> package, unfortunately, it had <a href="https://github.com/androidx/media/issues/1642">the same issue</a>.</p><p>Beyond just this assertion, the architecture of AdsLoader and AdPlaybackState was built around a &quot;Shared State&quot; model. In a multi-period timeline (e.g., <em>Period 0: Logo</em> -&gt; <em>Period 1: Movie</em>), Exoplayer applied the <strong>same</strong> AdPlaybackState to every period.</p><p>Because of this shared state we couldn’t just remove the assertion and move on with life, there were problems after that like a Preroll scheduled at 0s would try to play at the start of <em>every</em> period (Logo start, Movie start etc.), or Preroll will not play at ingress points like Continue watching from app.</p><h3><strong>Key Findings</strong></h3><p>ExoPlayer’s native multi-period DASH support has no client-side ad insertion. assert(periodCount == 1) assertions crash playback outright on multi-period streams (<a href="#">related discussion</a>).</p><p>Four problems followed from that root constraint:</p><p><strong>AdPlaybackState</strong> is designed for single-period timelines — applied to multi-period content, ads repeat across periods.</p><p><strong>Cue-point alignment</strong> breaks at period boundaries — ad breaks trigger early, late, or not at all.</p><p><strong>Preroll handling</strong> requires explicit edge-case logic: play once on cold start, suppress on re-entry from Continue Watching and equivalent ingress points.</p><p><strong>Backward compatibility</strong> forced a split serving strategy — single-period DASH for older clients, multi-period for newer ones.</p><p>Extending AdsMediaSource to handle this required architectural changes to ExoPlayer&#39;s ad handling layer. Rollout was phased, with playback failure rate, buffer times, and ad impressions as the primary watch metrics.</p><h4>Insights</h4><p>These findings underscored the need for a flexible, scalable approach to ad insertion in multi-period DASH content. Addressing ExoPlayer’s limitations and rethinking AdPlaybackState and cue-point handling laid the groundwork for a solution balancing technical feasibility and user experience.</p><h3>Methodology</h3><p>To address the challenges of enabling ads on multi-period DASH content in ExoPlayer, the team adopted a systematic and iterative approach:</p><ul><li>Audited ExoPlayer’s DASH manifest handling and ad timeline management to locate all single-period assumptions.</li><li>Prototyped assertion removal on multi-period manifests; used observed failures to surface edge cases early.</li><li>Built custom AdPlaybackState logic to split the main state into period-specific instances, scoping each ad to its designated period.</li><li>Implemented cue-point realignment relative to each period’s start time.</li><li>Replaced SinglePeriodAdTimeline with a custom MultiPeriodAdTimeline.</li><li>Gated multi-period DASH behind a version check; legacy clients continue on single-period.</li><li>Tested preroll/midroll, cross-period seeking, and platform variants (Android, Android TV, Fire TV).</li><li>Phased rollout starting with select content; monitored failure rates, buffer times, and ad impressions before expanding.</li></ul><h3>The Solution: A Custom MultiPeriodAdTimeline</h3><p>Two quick options were ruled out early — removing the assertion and upgrading to Media3 (covered in root cause).</p><p>Three turnkey approaches were evaluated:</p><ul><li>Separate players for ads and content</li><li>Client-side playlist stitching of single-period content and ad clips</li><li>ClippingMediaSource + ConcatenatingMediaSource to simulate a single timeline</li></ul><p>All three introduce buffering at player or content switches, lose AdsMediaSource capabilities (timeline management, seek handling), and scatter implementation complexity across multiple teams.</p><h4>Solution : MultiPeriodAdTimeline</h4><p>The chosen path: replace SinglePeriodAdTimeline with a custom MultiPeriodAdTimeline.</p><p><strong>First attempt</strong> — apply AdPlaybackState only to the content period; treat disclaimer and credits as ad-free.</p><p>Simple to implement. Two problems:</p><ul><li>Pre-roll plays after the disclaimer, not at stream start</li><li>No ads in the credits period — breaks standard behavior where the last cue-point fires on a direct seek to end</li></ul><p><strong>Second attempt</strong> — create dedicated AdPlaybackState per period (logo, content, dub card), each including a pre-roll to handle Continue Watching entry points.</p><p>Works, but hardcodes period sequence and count on the client. Any structural change to the stream breaks compatibility. Streaming team loses the freedom to modify period structure independently.</p><h4>Breaking the Monolith</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Oh2XgcxQ1xAMxcmIRNkyCg.png" /><figcaption>Propagate Playback State to Each Period</figcaption></figure><p>We took a step back and tried to generalize the solution. The biggest challenge was the “Shared State.” We handled all periods equally. We still have single main AdPlaybackState with cue-points as if we had single period content.</p><p>Internally its duplicated and <strong><em>transformed</em></strong> <strong><em>specially for each period. </em></strong>Any modifications happen on the main AdPlaybackState and are mirrored internally when it is updated.</p><p>We decided to keep indexing of ad breaks in AdPlaybackStates for each period the same. In each period will be all the cue-points in the same indices, but some will be ignored (skipped). Each period will have this transformations for each cue-point.</p><ul><li>Subtract start position of each period — times are relative to period start — some cue-points may end up being negative, but that is fine for us.</li><li>Mark cue-points after the period end as skipped. These will be played in the following periods</li></ul><p>That is it. A lot of investigation and iterations crystalized in few simple rules. After the transformation, the cue-point positions were as intended and all ad breaks are triggered even across period boundaries. From user experience, there is no difference between multi period and single period content. In the end <strong><em>single period is only a special case of multi period content now.</em></strong></p><h4>The Indexing Problem</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Oxn9qfqQbVsqukhztY4yOw.png" /><figcaption>Constant Break indices in Each Period</figcaption></figure><p>ExoPlayer identifies ad groups by <strong>index</strong>. If we simply removed “non-relevant” ad groups from a period’s state, the indices would shift, causing the player to play the wrong ad or crash.</p><p><strong>The Fix:</strong> We kept the Ad Group count constant across all periods.</p><ul><li>If Ad Group #1 belongs to Period 1, but we are currently configuring Period 0, we mark Ad Group #1 as SKIPPED in Period 0&#39;s state using AdPlaybackState.withSkippedAdGroup().</li><li>This ensures adGroupIndex 1 always refers to the same logical ad break, regardless of which period is currently active.</li></ul><h4>Handling Continuity</h4><p>We had to ensure that seeking across period boundaries didn’t re-trigger ads. Using the global bookmarking map, if a user watches an ad in Period 1 and seeks back to Period 0, the player knows that logic “Ad Break #X” is already played.</p><h4>Backward Compatibility</h4><p>We served multi-period DASH only to newer app versions; legacy versions continued with single-period DASH.</p><h3>Summary</h3><p>The implementation of ads on multi-period DASH content in ExoPlayer yielded the following outcomes:</p><h4>Technical Achievements</h4><ul><li><strong>Seamless Ad Playback</strong>: Ads are now played at the correct times, without repetition, even across period boundaries.</li><li><strong>Accurate Cue-point Handling</strong>: Dynamic adjustment of cue points ensures ads trigger precisely as intended.</li><li><strong>Robust Backward Compatibility</strong>: Users on older app versions experience no disruption, as they continue to receive single-period DASH.</li><li><strong>Performance Metrics</strong>: Key metrics such as start lag, buffering, ad impressions, and playback failure rates remained within acceptable thresholds throughout the rollout.</li></ul><h4>User Experience</h4><ul><li><strong>Dynamic Content Delivery</strong>: The platform can now insert disclaimers, dub cards, and localized content dynamically, enhancing personalization.</li><li><strong>Uninterrupted Viewing</strong>: Users experience smooth transitions between content and ads, with no playback disruptions.</li></ul><h4>Business Impact</h4><ul><li><strong>Uninterrupted Monetization</strong>: No impact on monetization as we modernized our media stack.</li><li><strong>Operational Efficiency</strong>: Reduced need for re-encoding and streamlined content pipelines have lowered operational overhead.</li><li><strong>Cost Savings: </strong>Re-encoding a large media library like JioHotstar’s would have been cost-prohibitive. With this solution, no separate encoding is needed for existing content.</li></ul><h3>Conclusion</h3><p>“Simple” features like adding a 5-second logo or disclaimer often hide iceberg-sized engineering challenges. By diving deep into the internals of ExoPlayer and rethinking how AdPlaybackStates are managed, we turned a hard constraint into a flexible capability.</p><p>If you are facing a similar issue, we have proposed these changes to be merged in upstream on <a href="https://github.com/androidx/media/pull/2501">Media3 Github repo here</a>.</p><p>This project reinforced a key lesson for us: sometimes the best way to move forward isn’t to work <em>around</em> the platform (multi-player), but to improve the platform itself!</p><p><em>Want to dig into player internals and contribute back to projects like ExoPlayer? Do check out </em><a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering"><em>open roles</em></a><em> if you want to build for millions of customers, who will use features that you build!</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=db8e8c00e521" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/monetisation-multi-period-dash-x-exoplayer-db8e8c00e521">Monetisation : Multi-period DASH x ExoPlayer</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Gen AI Video — Building Scalable Validation Framework]]></title>
            <link>https://blog.hotstar.com/building-scalable-validation-framework-for-video-generation-6c67d1177ce2?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/6c67d1177ce2</guid>
            <category><![CDATA[ai-video-generation]]></category>
            <category><![CDATA[generative-ai-use-cases]]></category>
            <category><![CDATA[ai-validation]]></category>
            <dc:creator><![CDATA[Sagar Tekwani]]></dc:creator>
            <pubDate>Wed, 18 Feb 2026 03:42:23 GMT</pubDate>
            <atom:updated>2026-02-18T03:42:22.487Z</atom:updated>
            <content:encoded><![CDATA[<h3><strong>Gen AI Video — Building A Scalable Validation Framework</strong></h3><blockquote>We discuss building strong eval frameworks as part of our Generative AI studio to ensure that generated video maintains a high quality bar.</blockquote><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*hOF71rV-cfi6sJwXo9tV8A.png" /></figure><h3>Introduction</h3><p>As JioHotstar scales to serve one of the world’s largest streaming audiences, generative AI(GenAI) offers a powerful opportunity to accelerate content creation, adapt stories across languages and regions, and unlock creative workflows that traditional pipelines cannot match.</p><p>Customers evaluate AI-generated video with the same standards they apply to premium productions; any drift in character identity, structural deformity, abrupt scene shift, or unsafe visual undermines realism instantly.</p><p>Generative models, being probabilistic, introduce such inconsistencies naturally, and at our scale even rare defects accumulate into meaningful quality gaps. Ensuring stable characters, coherent locations, and safe content therefore becomes a scientific challenge central to customer acceptance.</p><p>To address this, we built a <strong>Validation Framework</strong> that operates as a first-class component of the generation pipeline. This closed-loop system enforces quality, safety, and consistency at the same cadence as creation, enabling generative video to meet production-grade expectations on the JioHotstar platform.</p><h3>System Overview: The Validation Layer</h3><p>The video generation process starts with editorial scripts, from which the system extracts characters, locations, accessories, and context. These entities drive keyframe generation, which expands into video clips and assembles into the final video.</p><p>The <strong>Validation Framework</strong> (Fig. 1) spans all stages of this process and operates synchronously within the workflow. It evaluates intermediate outputs, flags issues early, and triggers targeted regeneration or parameter adjustments to maintain quality. Brand detection and safety checks act as hard gates, while other modules guide regeneration to preserve consistency and visual integrity.</p><p>Each module governs a specific quality dimension, character consistency, deformity, scene continuity, brand safety, content appropriateness, or story coherence and emits go/no-go signals that control progression. The system logs all validation outcomes and samples them for periodic human review to support ongoing calibration.</p><p>Together, these modules form the framework’s <strong>control surface</strong>, defining the quality of generative video. Most components operate at production readiness, while long-range story and concept continuity remains an active area of experimentation as we continue refining metrics and validation strategies.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*PiPJW32Mpdhx1ntQK7Ho3Q.png" /><figcaption><em>Fig.1: Overview of Developed Validation Framework</em></figcaption></figure><h3>Character Consistency</h3><p>Character consistency ensures that a character’s <strong>visual identity</strong> remains stable across all frames and scenes. Generative models, being stochastic, can drift subtly in facial features, proportions. These deviations break temporal realism and make the sequence unusable for production.</p><p>To quantify consistency, we represent each character through frame-level embeddings and compare them against a reference “hero character” image approved by the creative team. This hero image serves as the canonical visual anchor for that character.</p><p>We use an ensemble of independent facial similarity models (e.g., <em>Buffalo-L</em>, <em>Antelope-v2</em>, <em>FaceNet, etc.</em>), each fine-tuned for <strong>intra-character identity matching</strong>. Rather than representing a face with a single global embedding, these models extract <strong>multiple localised facial feature vectors</strong> corresponding to stable semantic regions of the face (such as eyes, nose bridge, jawline, and facial contours).</p><p>Sampling multiple localized descriptors improves robustness to pose changes, partial occlusion, lighting variation, and expression drift failure modes that are common in video generation but underrepresented in still-image similarity tasks. Each character, <em>k,</em> is represented in form of embeddings as f_hero^(k)</p><p>For a given frame <em>i</em>, each model m​ produces an embedding f_i^(k,m)​. We compute similarity using <strong>cosine similarity</strong>, which measures angular alignment in embedding space and remains invariant to feature magnitude.</p><p>This property is critical because generative models can alter contrast, illumination, and texture intensity without changing identity. <strong>Distance-based metrics</strong> (Euclidean or Manhattan) distance are sensitive to these magnitude shifts and empirically produce unstable thresholds across frames.</p><p>We also experimented with <strong>Jaccard similarity</strong> on facial embeddings but observed weaker alignment with <strong>human-in-loop evaluations</strong>. The similarity between the generated frame for a character <em>k</em> and its corresponding hero image for a model <em>m</em>, <em>S_i^(k,m)</em> is computed as:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/554/1*1rmvsE3Qb1CTcZs-OCD6pg.png" /></figure><p>Each model has a threshold m calibrated on <strong>human-labeled data.</strong> The binary decision per model m for each character k is:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/626/1*tQvFqPVlDG9_mbzpTqL_vA.png" /></figure><p>The final consistency decision <em>C</em> uses an <strong>ensemble aggregation </strong>across all models (N) and characters, tuned for <strong>recall maximization</strong>:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/554/1*nPG9dNM59dlSQ1JJY5Z90g.png" /></figure><p>where, m​ represents each model’s reliability weight and defines ensemble sensitivity.</p><p>This design ensures that every potential inconsistency is captured, even if one model fails, prioritizing recall over precision. Minor false positives are filtered through, ensuring consistency over short frame windows rather than isolated detections. Thresholds m and m are periodically re-calibrated using <strong>human-in-loop feedback</strong>. Annotators review flagged segments, refine labels, and feed corrections back into the validation loop. Thus, by fusing diverse fine-tuned models and anchoring every comparison to the hero reference, the system maintains stable character identity across the generation pipeline, regardless of lighting or scene transition.</p><p>For, side-profile consistency, which often reveals identity drift missed in frontal views. We generate strict left and right 90-degree profile references from the approved hero image, excluding frontal or angled views. We store the <strong>front, left, and right profiles</strong> as canonical anchors and compare generated frames against them to validate identity across viewing angles. This check surfaces profile-specific inconsistencies early and triggers regeneration when needed.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*T8jH7gdKIhEtF8YkVHimYA.png" /><figcaption><em>Fig.2: Working of Character Consistency Framework</em></figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*2PeHKnBwtismAH-63JfJXA.png" /><figcaption><em>Table-1: Comparison of Hero Image with different generations</em></figcaption></figure><h3>Character Deformity Detection</h3><p>Character deformities break visual realism immediately. Generative models can produce warped limbs, misaligned joints, or anatomically implausible body structures when spatial constraints fail during sampling. To detect these failures reliably, we built a deformity-classification pipeline grounded in curated abnormality data and a trained YOLO-family detector.</p><p>We first evaluated general-purpose multimodal models such as <strong>Gemini</strong> and <strong>Qwen-VL</strong> on deformity detection. These models achieved only <strong>≈40% recall</strong> on our internal human-annotated deformity dataset, and they consistently failed to detect subtle or multi-region structural distortions. <strong>This baseline confirmed that deformity detection requires a dedicated model trained on explicit abnormality signals. </strong>To detect anatomical distortions reliably, we built a deformity-detection module optimized for high recall and early intervention during video generation.</p><ul><li>Used Tencent’s Distortion dataset (<em>Predicting Distortion in Real-World Human Images</em>) as the base corpus and curated it using human-in-loop review to improve label reliability.</li><li>Applied a segmentation-guided cleaning pipeline to remove annotations outside the human region, discard low-confidence samples pseg​, and filter out deformity regions below a minimum area threshold Amin​; segmentation masks also provided body-part bounding boxes.</li><li>Trained a YOLO-family detector on the curated dataset to localize and classify deformities across full-body and body-part crops, explicitly optimizing for high recall., Achieving ~<strong>35% recall lift</strong> on the same human-annotated evaluation set in comparison to Gemini and Qwen.</li><li>Integrated the detector across all generation stages to flag anatomical failures early and trigger regeneration or parameter adjustments before outputs propagate downstream.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ixDPPIZNDvQCuWGOQxPkpQ.png" /><figcaption><em>Fig.3: Deformity Detection Framework for limb abnormalities</em></figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*FCJkhR6QAUNjqrXaAKgXYQ.png" /><figcaption><em>Fig.4: Deformity Detection Framework for facial abnormalities</em></figcaption></figure><h3>Location Consistency</h3><p>Location consistency ensures that generated scenes remain faithful to the <strong>scripted description</strong> and stable across <strong>multiple scenes within the same location</strong>. This is extremely important since drift in layout, lighting, spatial structure, or persistent objects breaks continuity and degrades perceived quality.</p><p><strong>Consistency with scripted location and object constraints: </strong>During script parsing, LLMs extract structured location descriptions that capture spatial layout, environmental cues, lighting intent, and object-level constraints. These descriptions define both the expected set of objects and a subset of <strong>mandatory elements</strong> whose presence must be preserved.</p><p>We evaluate <strong>object presence</strong> using vision language models (VLMs) and flag in case if required objects are missing. In parallel, we assess overall scene fidelity by aligning text embeddings derived from the location description with visual scene embeddings extracted from generated keyframes using our VLM stack (e.g., Gemini, Qwen).</p><p>We calibrate <strong>similarity thresholds</strong> through human-in-loop (HIL) evaluation, selecting cutoffs that best correlate with human judgments of scene correctness. Frames that fall below the calibrated threshold indicate semantic or structural violations and trigger regeneration or parameter adjustment. Figure 5 illustrates how alignment scores reflect adherence to scripted location and object constraints.</p><blockquote>Scene Visualization:<strong> </strong>Interior. Police interrogation room. Dim overhead lighting. Metal table bolted to the floor. Two chairs facing each other. One-way mirror. No windows.</blockquote><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xiUH_ZdnlC7nIAqN-q3dVw.png" /><figcaption><em>Fig.5: Location Consistency with Script&lt;&gt;Object alignment</em></figcaption></figure><p><strong>Consistency across scenes within the same location: </strong>For locations that recur across multiple scenes, we treat the first validated keyframe as the one achieving the highest text–image alignment score as the <strong>location anchor</strong>. Using our vision–language model (VLM) stack, we extract scene-level embeddings from subsequent frames and compare them against this anchor via cosine similarity to detect structural and stylistic drift. This formulation allows us to enforce consistency in spatial layout, lighting characteristics, and persistent background elements when the same room, street, or set reappears at different points in the video. <strong>Figure 6</strong> illustrates this anchor-based consistency check.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*x0U1MUP-25NjmyVZsko32Q.png" /><figcaption><em>Fig.6: Location Consistency across scenes</em></figcaption></figure><p>LLMs provide strong priors when generating detailed location descriptions, but they do not yet reliably detect subtle spatial or geometric inconsistencies in generated frames. We are therefore exploring additional approaches, including embedding-based scene classifiers, layout-consistency models, and structure-aware validators, to strengthen this module. These efforts aim to convert location consistency into a fully measurable and enforceable dimension within the validation framework.</p><h3>Brand Logo Detection/Safety Checks</h3><p>Brand-logo/Safety violations act as <strong>hard safety gates</strong> in the generation pipeline. The system blocks any output containing such elements and triggers regeneration until the output passes all safety checks. The validation loop operates as follows: the detector scans each generated unit, flags any violation, the system regenerates the content with adjusted constraints, and the detector re-evaluates the regenerated output. If repeated attempts fail, the system escalates the case for manual review. This loop ensures no flagged safety issue propagates downstream.</p><p>We enforce these safety dimensions through <strong>prompt-level controls</strong> and <strong>automated detection</strong>. During script parsing, LLMs generate structured content descriptions, including required or disallowed visual categories. These descriptions guide the image-generation model to avoid branded items and unsafe content.</p><p>We curated datasets for both tasks: brand/logo samples across categories such as laptops, consumer electronics, apparel, etc. and safety violations across violence, child-abuse, racism, hate symbols, and other violation types.</p><p>Using these datasets, we optimized prompts and detection thresholds for <strong>high recall</strong>, achieving <strong>&gt;95% recall</strong> on internal evaluations. When the detector identifies a brand or NSFW instance, the system regenerates the output with stricter constraints to remove the violation.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*XZZnbMrU2UUCdExv6t7bWw.png" /><figcaption><em>Fig.7: Brand Logo Detection across categories</em></figcaption></figure><h4>Engineering Controls to Prevent Recurrence of Violations</h4><p><strong>Context-Aware Decoding:<br></strong> We add structured negative constraints that suppress brand names, logos, or unsafe categories during generation. These constraints adjust the decoding trajectory of the image-generation model and reduce the probability of producing forbidden visual elements.</p><p><strong>Adaptive Prompt Rewriting:<br></strong> When a violation is detected, the system rewrites the prompt by tightening constraints, clarifying allowed content, and removing ambiguous phrasing. These rewritten prompts condition the regeneration step and help eliminate repeated violations across attempts.</p><h3>Summary and Future Directions</h3><p>This work presents a <strong>validation framework for generative video</strong> that operates as a first-class component of the generation pipeline. By integrating validation directly into the workflow, the system detects and corrects character inconsistency, anatomical deformities, location drift, and safety violations during generation.</p><p>The framework applies recall-first validation, multi-model similarity checks, vision–language alignment, and thresholds calibrated through human feedback to convert qualitative notions of visual quality into enforceable signals. This design prevents narrative-breaking errors from propagating while allowing controlled creative variation.</p><p>Next, we will formalize a unified <strong>evaluation and metrics layer</strong> that measures video, audio, lip-sync, and temporal consistency and supports systematic optimization.</p><p>We will also extend the framework to address <strong>long-range scene and concept continuity</strong>, enforcing coherence across scenes, episodes, and story arcs. These extensions will complete the quality stack required to deploy generative video systems at production scale.</p><p><em>Want to be at the forefront of generative video in India? Do check out </em><a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering"><em>open roles</em></a><em> if you want to build for millions of customers, who will use features that you build!</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=6c67d1177ce2" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/building-scalable-validation-framework-for-video-generation-6c67d1177ce2">Gen AI Video — Building Scalable Validation Framework</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Modernizing Dependency Management: Beyond CocoaPods (Part 2 — The Execution)]]></title>
            <link>https://blog.hotstar.com/modernizing-dependency-management-beyond-cocoapods-part-2-the-execution-bf3b9d739efb?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/bf3b9d739efb</guid>
            <category><![CDATA[swift]]></category>
            <category><![CDATA[mobile-development]]></category>
            <category><![CDATA[ios-development]]></category>
            <category><![CDATA[swift-package-manager]]></category>
            <category><![CDATA[dependency-management]]></category>
            <dc:creator><![CDATA[Saurabh Kapoor]]></dc:creator>
            <pubDate>Fri, 13 Feb 2026 04:26:22 GMT</pubDate>
            <atom:updated>2026-02-13T04:26:20.830Z</atom:updated>
            <content:encoded><![CDATA[<h3>Modernizing Dependency Management: Beyond CocoaPods (Part 2 — The Execution)</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1022/1*uo1p1ZKkrCqxXMBsbm6uxw.jpeg" /></figure><h3>Recap</h3><p>In <a href="https://blog.hotstar.com/dependency-management-our-journey-beyond-cocoapods-part-1-the-strategy-c3c7874566a9">Part 1</a>, we shared the strategy — why we had to move away from CocoaPods, how we evaluated our options, and the phased migration plan we designed to modernize 60 pods over 2 quarters without disrupting feature development.</p><p>We bring it all home and share our learnings and scars.</p><h3>Phase 1: Proving That SPM Could Coexist</h3><p>Big bang approaches for a product that supports over a billion Indians is not on the menu. We could not stop the train and we had to upgrade while we ran as fast as we were.</p><blockquote><em>Swift Package Manager (SPM) had to live alongside CocoaPods without disrupting day-to-day development.</em></blockquote><p>This was not negotiable. It was also about building confidence gradually.</p><p>Our dependency graph was already complex, and CocoaPods was deeply embedded in how the app was built, tested, and shipped. Introducing a second dependency manager into that ecosystem was risky. If coexistence failed, everything downstream would be compromised.</p><p>So we treated Phase 1 as a controlled experiment.</p><h4>Defining Success</h4><p>We were explicit about what success looked like:</p><ul><li>Developers should not need to change their workflows</li><li>CI pipelines must remain stable</li><li>No runtime regressions</li><li>CocoaPods must remain fully functional</li></ul><h4>Choosing the Right First Dependencies</h4><p>We intentionally avoided internal SDKs in this phase. Instead, we chose third-party libraries that already had mature SPM support and met three criteria: widely used in the app, minimal customization, and no deep runtime coupling with internal frameworks or development pods.</p><p>This allowed us to isolate SPM behavior without risking business-critical flows. By migrating only a handful of carefully selected dependencies, we could observe real-world behavior without destabilizing the system.</p><h4>Invisible to Developers</h4><p>One of the most important constraints we enforced was developer invisibility.</p><p>Developers continued to open the same workspace, build using the same schemes, run tests the same way, and rely on the same CI signals. There were no new scripts to run, no new commands to remember, no changes to onboarding docs.</p><p>SPM dependencies resolved automatically by Xcode in the background. If someone hadn’t been told we were testing SPM, they wouldn’t have noticed.</p><h4>Phase 2: Internal Pods and Binary Distribution</h4><p>Phase 1 inspired confidence and we ramped up and doubled down. Internal pods were where things got interesting.</p><p>These weren’t isolated third-party libraries. They were shared across multiple apps, some across Android, and were actively developed. Moving them required more than a format change — it required rethinking how we distribute internal code.</p><h4>The Promise of Binary Targets</h4><p>Moving our internal pods to SPM binary targets felt like a natural evolution. Using .xcframework with .binaryTarget allowed us to distribute prebuilt artifacts instead of rebuilding large internal modules every time.</p><pre>.binaryTarget(<br>    name: &quot;CoreSDK&quot;,<br>    url: &quot;https://example.com/CoreSDK.xcframework.zip&quot;,<br>    checksum: &quot;...&quot;<br>)</pre><p>This allowed us to preserve encapsulation, reduce build times, and decouple SDK evolution from app builds.</p><p>But this phase also exposed one of our biggest challenges.</p><h4>The Problem: Binary Targets and Private Repositories</h4><p>Very quickly, we ran into a problem that wasn’t obvious from the documentation.</p><p>Swift Package Manager assumes that binary artifacts are publicly accessible. When SPM encounters a binaryTarget(url:), it attempts to download the zip file, verifies the checksum, and caches the artifact locally. What it does <em>not</em> do is authenticate.</p><p>That assumption works fine for open-source packages hosted publicly — but completely breaks down for private internal SDKs.</p><p>We hosted our .xcframework.zip files as GitHub release assets inside private repos. Everything looked correct: URL was valid, checksum matched, artifact was present. Yet builds consistently failed:</p><pre>Failed to download binary artifact<br>The requested URL returned error: 404</pre><p>The file existed. The problem was subtle but critical: SPM was making unauthenticated HTTP requests to private URLs. GitHub correctly responded with 404, not 401, masking the real issue.</p><h4>Why This Was a Big Deal</h4><p>This wasn’t just an inconvenience — it had architectural implications:</p><ul><li>Developers couldn’t resolve packages locally</li><li>CI pipelines failed deterministically</li><li>Binary targets became unusable for internal SDKs</li><li>The entire binary-distribution strategy was at risk</li></ul><p>At this point, we had to pause and ask: <strong><em>Can SPM actually work for private, enterprise-scale binary distribution?</em></strong></p><h4>Exploring Workarounds</h4><p>We explored multiple approaches, each with trade-offs:</p><p><strong>Making repositories public</strong> — Immediately ruled out. Internal SDKs contain proprietary logic.</p><p><strong>Embedding tokens in URLs</strong> — Technically possible, but unacceptable. Security risk, tokens leak via logs, impossible to rotate safely.</p><p><strong>Git LFS</strong> — Workable, but introduced large repository sizes, slower clones, and additional tooling overhead. Didn’t scale well for frequent SDK releases.</p><p><strong>Artifact repositories (S3/Nexus)</strong> — Viable, but required additional infrastructure, credential management, and URL signing logic.</p><p>We wanted something simpler for GitHub-hosted binaries.</p><h4>The Solve: .netrc Authentication</h4><p>The solution came from understanding how SPM downloads binaries.</p><p>SPM relies on standard system networking under the hood. That means it respects .netrc credentials, just like curl or git. By configuring authentication at the system level, we could allow SPM to fetch private binaries without changing a single line of Package.swift.</p><pre># ~/.netrc<br>machine github.com<br>login GITHUB_USERNAME<br>password GITHUB_PERSONAL_ACCESS_TOKEN</pre><p>Once this file was present, SPM successfully authenticated, binary artifacts downloaded correctly, checksums validated as expected, and builds became deterministic again.</p><p>Most importantly, this worked identically on developer machines and CI runners.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*p9_PIC6RcQuH075qlgLOhA.png" /></figure><h4>Hardening : Making It CI-Friendly</h4><p>On CI, we injected the .netrc file securely at runtime using secrets:</p><pre>echo &quot;machine github.com login $GITHUB_USER password $GITHUB_TOKEN&quot; &gt;&gt; ~/.netrc<br>chmod 600 ~/.netrc</pre><p>This gave us no credentials in source control, easy token rotation, and clear audit boundaries. It also aligned well with GitHub Actions and self-hosted runners.</p><h4>What This Unlocked</h4><p>Solving this problem unlocked the full potential of SPM binary targets:</p><ul><li>Internal SDKs could be versioned independently</li><li>App builds became significantly faster</li><li>SDK releases became predictable artifacts</li><li>Teams consumed binaries without worrying about source-level coupling</li></ul><p>Onwards!</p><h3>Phase 3: Development Pods → Swift Packages</h3><p>Development pods were not just dependencies — they were living parts of the app, evolving alongside features, touched daily by multiple teams. They were also deeply intertwined with how our codebase was structured.</p><h4>The Challenge: Living Code, Not Just Dependencies</h4><p>Our development pods served a very specific purpose: they allowed teams to iterate on shared modules without releasing binaries, supported rapid local changes and debugging, and encoded architectural boundaries inside the Podfile.</p><p>Over time, they became like an extension of the app — not just dependencies.</p><p>Replacing them with Swift packages meant answering a hard question: <em>Can we retain the same developer experience without CocoaPods doing the heavy lifting for us?</em></p><h4>Preserving the Existing Structure</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*9eV0SZYqS6nqBfdMlUxT-w.png" /></figure><p>One of our biggest concerns was accidentally reshaping the codebase. We did not want to flatten modules, merge responsibilities, or introduce artificial package boundaries.</p><p>Instead, we followed a strict rule: <strong>every development pod becomes a Swift package with the same conceptual boundaries.</strong></p><p>That meant one pod became one package, with the same folder structure, same ownership, and same responsibility. This discipline paid off later when debugging regressions and onboarding developers.</p><h4>Language Boundaries: Objective-C and Swift</h4><p>Swift Package Manager supports Objective-C, but it does not allow multiple languages within the same target. A single target cannot contain both Swift and Objective-C sources.</p><p>Several of our development pods relied on exactly that — Swift and Objective-C files coexisting within the same logical module, with bridging handled implicitly by the build system. Under SPM, this was no longer possible.</p><p>To move forward, we restructured these modules intentionally. Objective-C code was moved into dedicated package targets with explicitly defined public headers. Swift targets then depended on these Objective-C targets through clear, declared dependencies.</p><p>This refactoring made language boundaries explicit and removed implicit bridging behavior. While it required effort, it resulted in a cleaner dependency graph and clearer ownership across modules.</p><h4>Platform Boundaries: XIBs and Multi-Platform Packages</h4><p>Legacy XML Interface Builder (XIB)s introduced a different kind of challenge — one tied closely to how Swift Package Manager treats multi-platform packages.</p><p>Under CocoaPods, platform-specific behavior could be handled through Podfile logic or build configurations. This allowed development pods to bundle UI resources that behaved differently across iOS and tvOS without that distinction being explicit in the code.</p><p>Swift Package Manager takes a stricter approach. Packages are inherently multi-platform, and resources declared in a package are shared across all supported platforms. There is no native way to conditionally include or exclude resources based on platform within the same target.</p><p>Several development pods contained XIBs that were platform-specific and loaded conditionally at runtime. Once these pods became Swift packages, those assumptions no longer held.</p><p>Rather than fragmenting packages or introducing complex runtime branching, we made a deliberate architectural decision: <strong>we moved most XIB-based UI into code.</strong></p><p>This shift reduced reliance on resource bundling, eliminated fragile platform assumptions, and aligned better with the multi-platform model encouraged by Swift Package Manager — even though it required more upfront refactoring.</p><h4>Build Behavior: Replacing Post-Install Scripts</h4><p>One aspect we underestimated initially was how much logic lived outside the code itself.</p><p>Over the years, our CocoaPods setup had accumulated substantial scripting inside the post_install block. These scripts handled modifying build settings across targets, injecting compiler and linker flags, adjusting deployment targets, patching generated project settings, code generation for modules, and handling branding assets.</p><p>CocoaPods made this convenient because it centralized these changes in one place. But once we moved away from CocoaPods, that implicit behavior disappeared immediately.</p><p>Swift Package Manager does not offer an equivalent of a post_install hook. That forced us to confront an important reality: <strong>a lot of critical build behavior was hidden in scripts that developers rarely looked at.</strong></p><p>To preserve correctness without reintroducing global magic, we deliberately moved this logic closer to where it actually mattered. Most essential scripting was redistributed into explicit Xcode build phases, scoped to the relevant app or framework targets. In some cases, we replaced scripts entirely by fixing the underlying configuration rather than patching it at build time.</p><p>This shift had two important effects: build behavior became more visible and discoverable, and changes were scoped to specific targets instead of being applied globally.</p><h4>Resources, Flags, and Conditional Logic</h4><p><strong>Resources</strong> — Assets that were automatically bundled by CocoaPods now had to be declared explicitly:</p><pre>.target(<br>    name: &quot;UserProfileKit&quot;,<br>    resources: [<br>        .process(&quot;Resources&quot;)<br>    ]<br>)</pre><p>This forced us to audit every resource and validate runtime access paths — something CocoaPods had silently handled for years.</p><p><strong>Build Settings</strong> — CocoaPods’ pod_target_xcconfig allowed us to inject compiler and linker settings easily. SPM requires these to be expressed explicitly:</p><pre>swiftSettings: [<br>    .define(&quot;ENABLE_LOGGING&quot;, .when(configuration: .debug))<br>]</pre><p>For edge cases, we used .unsafeFlags — but sparingly and deliberately. This made us more intentional about what each module actually required.</p><p><strong>Conditional Inclusion</strong> — In CocoaPods, it was common to conditionally include pods based on build configurations. SPM does not support conditional dependencies for custom build configurations.</p><p>This forced a shift in thinking. Instead of conditionally including dependencies, we moved toward conditionally <em>using</em>them via compile-time flags and feature gates:</p><pre>#if ENABLE_EXPERIMENTAL_FEATURE<br>// Feature-specific code<br>#endif</pre><p>This change improved clarity, even though it required refactoring.</p><h4>Keeping Development Fast</h4><p>A common fear with moving development pods to SPM is slower iteration. We paid close attention to this.</p><p>Local packages were referenced via relative paths, changes reflected immediately in the app, and the debugging experience remained intact. From a developer’s perspective, very little changed — which was exactly what we wanted.</p><h4>CI and Testing Implications</h4><p>Moving development pods affected CI in subtle ways. Test targets needed explicit dependency declarations, schemes had to be updated, and build order changed slightly.</p><p>This surfaced hidden assumptions in our pipelines — but fixing them made CI more robust and predictable.</p><h4>The Emotional Reality</h4><p>This phase took time. It required patience. And it touched a lot of code.</p><p>But it also marked a turning point. By the end of Phase 3, CocoaPods was no longer central to our architecture, Swift packages were no longer “new,” and the system felt cleaner and more explicit.</p><p><strong>We stopped thinking in terms of pods and started thinking in terms of modules with explicit contracts.</strong></p><h4>The Surprises We Didn’t See Coming</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1020/1*vBOg5RI_WkYuOI86y48vtg.jpeg" /></figure><p>We expected friction. We anticipated refactoring. What we didn’t fully anticipate was how many assumptions CocoaPods had been quietly absorbing for us over the years — assumptions that surfaced only once Swift Package Manager forced everything into the open.</p><p>These challenges didn’t appear all at once. They emerged gradually, often at inconvenient moments, and almost always in places we thought were already “done.”</p><h4>Duplicate Symbols at Runtime</h4><p>One of the more subtle issues we encountered was related to duplicate symbols — but not in the way they typically present themselves. These were not build-time or linker failures. Builds succeeded, targets launched normally, and at a glance everything appeared to be working.</p><p>The problems surfaced only at runtime.</p><p>Because Swift Package Manager builds packages as static libraries by default, the same dependency could end up being embedded multiple times through different dependency paths. This resulted in multiple instances of what was expected to be a single module being present in the process.</p><p>At runtime, this caused the wrong instance of certain symbols to be referenced:</p><ul><li>Global state would appear to reset or diverge</li><li>Dependency injection would resolve to unexpected instances</li><li>Mocks would not behave as expected</li></ul><p>In some cases, the only signal we had was a vague runtime warning:</p><pre>objc[12345]: Class MySharedService is implemented in both<br>/path/to/App.app/App and /path/to/AppTests.xctest/AppTests.<br>One of the two will be used. Which one is undefined.</pre><p>Nothing crashed. Nothing failed to launch. But from that point onward, behavior was undefined.</p><p>The challenge was not detecting the issue, but diagnosing it. Since nothing failed during build or launch, the failures initially looked like flaky tests or logical bugs rather than a dependency problem.</p><p>Resolving this required a careful audit of how dependencies were introduced across app, framework, and test targets. We had to ensure that shared modules were linked exactly once and that dependency graphs were consistent across targets.</p><p><strong>This was a strong reminder that test targets are not passive consumers of the app binary. They are independent bundles with their own runtime environment.</strong></p><h4>Builds That Succeeded but Crashed</h4><p>Some of the most difficult issues gave us the least amount of signal.</p><p>The app compiled successfully. There were no compiler errors. There were no linker warnings. And yet, the application crashed immediately at runtime.</p><p>These failures didn’t surface during build because, from the compiler’s perspective, everything was valid. All symbols were present, all dependencies resolved, and the binary was produced without complaint. The problem only emerged once the app launched and the runtime attempted to load and resolve those symbols.</p><pre>dyld: Symbol not found: _$s15MySharedModule16CriticalServiceC11sharedInstanceACvgZ<br>  Referenced from: /Applications/App.app/App<br>  Expected in: /Applications/App.app/Frameworks/MySharedModule.framework/MySharedModule</pre><pre>dyld: Library not loaded: @rpath/MySharedModule.framework/MySharedModule<br>  Reason: image not found</pre><p>Diagnosing these issues required stepping outside the usual compile–link–run mental model. We had to inspect the final app binary, verify which frameworks and libraries were actually embedded, and confirm that runtime search paths and linkage settings were aligned with how the dependencies were built.</p><h4>Saying Goodbye to Slather</h4><p>One of the bigger surprises had nothing to do with compilation, linking, or dependency resolution. It was code coverage.</p><p>For years, we had relied on Slather as our coverage tool. It was stable, familiar, and deeply integrated into our CI pipelines. We assumed it would continue to work as we moved dependencies to Swift Package Manager.</p><p>That assumption turned out to be wrong.</p><p>Slather does not support pure Swift packages as first-class citizens. It expects coverage to be generated from Xcode projects or workspaces, not standalone packages. As more of our codebase moved into Swift packages, coverage for package-based modules simply disappeared.</p><p>The only way to keep Slather working would have been to create artificial “host” projects whose sole purpose was to run package tests and collect coverage. That approach went directly against our goal of making systems simpler.</p><p>At that point, it became clear that the problem wasn’t Swift Package Manager — it was the tooling around it.</p><p>We switched to an xcresult-based coverage pipeline, parsing coverage directly from Xcode&#39;s native test result bundles. This aligned us with the direction Apple was already taking. Coverage became more accurate, easier to reason about, and independent of how the code was packaged.</p><h3>Phase 4: Removing CocoaPods</h3><p>By the time we reached this phase, CocoaPods was no longer doing much work.</p><p>All third-party dependencies had already moved to Swift Package Manager. Internal SDKs were consumed as binary targets. Development pods had been fully replaced with package-based modules.</p><p>And yet, CocoaPods was still there — quietly present, still wired into the system.</p><p>Removing it was less about technical effort and more about confidence.</p><h4>Knowing When We Were Ready</h4><p>For several weeks, CocoaPods existed in the repository almost as a safety net. During this period, we monitored build stability across all configurations and regions, CI performance and reliability, developer onboarding, and QA cycles without CocoaPods involvement.</p><p>Only after we had multiple successful releases — without touching CocoaPods at any stage — did we decide it was time to move on.</p><h4>The Final Delete</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*oHsBUFkMcLx6lZbkJqyAJw.jpeg" /></figure><p>The final step was straightforward: we removed the Podfile, the CocoaPods-generated .xcworkspace, and all remaining CocoaPods-related files and scripts.</p><p>There was no disruption, no rollback, and no follow-up fixes.</p><p>The system simply continued to work — without CocoaPods.</p><p><strong>The day we merged that PR felt like crossing a finish line we’d been racing toward for six months.</strong></p><h3>What This Journey Taught Us</h3><p>By the time CocoaPods was fully removed, the migration itself had stopped feeling like the most important outcome.</p><p>What mattered more was how the journey reshaped the way we think about our codebase, our tooling, and the systems that support day-to-day development.</p><p>This wasn’t just a dependency migration. It was a gradual recalibration of engineering discipline.</p><h4>Tooling Should Fade Into the Background</h4><p>The best tooling is the kind you don’t think about.</p><p>CocoaPods had accumulated years of scripts, configuration overrides, and implicit behavior. Swift Package Manager, in contrast, forced us to be explicit — but once configured, it largely disappeared into the background.</p><p>When developers no longer need to remember setup steps or debug dependency resolution, cognitive load drops. Productivity rises not because things move faster, but because there’s less to manage.</p><h4>Migration Is About Trust, Not Speed</h4><p>One of the most important decisions we made was not rushing.</p><p>By migrating in phases and allowing CocoaPods and SPM to coexist, we preserved trust: trust from developers that their workflows wouldn’t break, trust from QA that releases wouldn’t destabilize, trust from leadership that the migration wouldn’t impact delivery.</p><p>The time spent validating coexistence and waiting through release cycles was not overhead — it was risk mitigation.</p><h4>CI/CD Is Part of the Architecture</h4><p>Several challenges — especially around binary targets, coverage, and caching — forced us to acknowledge something we had underweighted before:</p><p>CI/CD is not infrastructure glue. It’s part of the system design.</p><p>Solving problems like authenticated binary downloads or deterministic package resolution required thinking beyond Xcode and into the pipeline itself. Once we did, CI became more reliable, more predictable, and easier to maintain.</p><h3>Closing Thoughts</h3><p>CocoaPods played a critical role in helping us scale our iOS and tvOS ecosystem at a time when the platform needed it most. It gave us structure, enabled modularization, and supported years of rapid development. For that, it deserves recognition.</p><p>Swift Package Manager represents where the Apple ecosystem is heading. It aligns more closely with Swift itself, integrates natively with Xcode, and encourages explicit, predictable dependency management. Adopting it wasn’t just a response to CocoaPods’ sunset — it was an investment in long-term maintainability and clarity.</p><p>The migration was not quick, and it was never meant to be. We approached it deliberately, prioritizing stability over speed and confidence over convenience. By moving in phases, preserving existing workflows, and validating each step through real release cycles, we ensured that modernization didn’t come at the cost of ongoing development.</p><p><strong>Six months. Two quarters. Sixty pods. Zero broken releases. One transformed dependency management system.</strong></p><p>If you’re standing at a similar crossroads, our advice is simple but hard-earned:</p><p><strong>Move with intent. Move carefully. And never break development in the process.</strong></p><p>We’re hiring for our client teams! Do check out <a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering">open roles</a> if you want to build it right, while you build for millions of customers, who will use features that you build!</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=bf3b9d739efb" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/modernizing-dependency-management-beyond-cocoapods-part-2-the-execution-bf3b9d739efb">Modernizing Dependency Management: Beyond CocoaPods (Part 2 — The Execution)</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Dependency Management: Our Journey Beyond CocoaPods (Part 1 — The Strategy)]]></title>
            <link>https://blog.hotstar.com/dependency-management-our-journey-beyond-cocoapods-part-1-the-strategy-c3c7874566a9?source=rss----dbc3fcbc7f07---4</link>
            <guid isPermaLink="false">https://medium.com/p/c3c7874566a9</guid>
            <category><![CDATA[ios-development]]></category>
            <category><![CDATA[swift-package-manager]]></category>
            <category><![CDATA[swift]]></category>
            <category><![CDATA[dependency-management]]></category>
            <category><![CDATA[mobile-app-development]]></category>
            <dc:creator><![CDATA[Saurabh Kapoor]]></dc:creator>
            <pubDate>Tue, 03 Feb 2026 08:58:51 GMT</pubDate>
            <atom:updated>2026-02-03T08:58:50.101Z</atom:updated>
            <content:encoded><![CDATA[<h3>Modernizing Dependency Management: Beyond CocoaPods (Part 1 — The Strategy)</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*LKzyVxVFyIXupSUYoQaH5g.jpeg" /></figure><blockquote>CocoaPods formed the spine of our iOS build workflow. In this two part blog, we’re sharing our journey to transparently switch away from CocoaPods to SPM, minus the drama, but plus all the scars!</blockquote><h3>Where We Started</h3><p>CocoaPods was at the heart of our development workflow — for everything. Dependency resolution, configuration overrides, CI integration — all became tightly coupled to CocoaPods. It stopped being a convenience layer and became part of our infrastructure.</p><p>Then came the <a href="https://blog.cocoapods.org/CocoaPods-Specs-Repo/">announcement</a> that changed everything.</p><blockquote><strong>CocoaPods’ trunk would become read-only in December 2026.</strong></blockquote><p>This wasn’t just an ecosystem update. It was a tectonic shift. A read-only trunk meant no new pod versions, no straightforward path to adopt upstream fixes, and increasing exposure to unpatched issues over time. Any disruption here wouldn’t just affect builds — it would impact active feature development, CI stability, and release confidence across teams.</p><p>So when the announcement landed, the question wasn’t <strong><em>“Should we move?”</em></strong> That decision had effectively been made for us.</p><p>What followed was one of the most ambitious infrastructure initiatives we’ve undertaken — a six-month effort, touching every corner of our codebase, and ultimately transforming how we build and ship our apps.</p><p>This was open-heart surgery on a system that powered millions of app sessions every day — performed while the patient was still running around</p><p>This is that story.</p><h3>Choices, choices..</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*7GFT2BmFXcc4bDk_evUDmw.jpeg" /></figure><p>Before committing to any solution, we took a deliberate step back and evaluated the dependency management landscape as a whole. Since we had to move, we wanted to make as much of a future proof choice as possible.</p><p>Several factors mattered deeply to us:</p><ul><li><strong>Scalability</strong> — Handle the size and complexity of our codebase, while remaining adaptable.</li><li><strong>Learning curve</strong> — It couldn’t slow teams down or force widespread workflow changes.</li><li><strong>Future-proofing</strong> — We wanted to align with the direction the Apple ecosystem was moving.</li><li><strong>Tooling integration</strong> — It had to work well with project generation and tooling we already relied on.</li></ul><p>With those constraints in mind, we evaluated the available options.</p><h3>Carthage</h3><p>Carthage is lightweight, decentralized, and intentionally avoids modifying Xcode projects — qualities that align well with engineering simplicity. For smaller projects or teams with straightforward dependency needs, it can be an excellent choice.</p><p>However, managing binary distribution at scale, supporting internal SDKs, handling complex build configurations, and maintaining consistent tvOS support required more manual orchestration than we were comfortable with. Over time, this would have shifted operational complexity from tooling into team workflows.</p><p>Carthage wasn’t a bad fit universally; it just wasn’t the right fit for our ecosystem.</p><h3>Bazel and Buck</h3><p>We also explored Bazel and Buck — not dependency managers in the traditional sense, but complete build systems. Both are powerful and proven at massive scale, offering deterministic builds, strong caching, and sophisticated dependency graphs.</p><p>However, adopting either would have meant replacing our entire build system, moving away from Xcode’s native build model, and introducing a steep learning curve for developers whose daily workflows are deeply tied to Xcode.</p><p>For our teams, the cost of that transition far outweighed the benefits.</p><h3>Swift Package Manager (SPM)</h3><p>Swift Package Manager wasn’t perfect. It lacked some of the flexibility we were used to, and certain features required workarounds. But it had two qualities that ultimately mattered most.</p><p>First, it was <strong>native</strong> — integrated directly with Xcode, aligned with Swift’s evolution, and benefiting from ongoing investment by Apple. Second, it allowed us to <strong>preserve our existing project structure</strong> while gradually modernizing it. We could migrate incrementally, validate changes in production, and avoid large-scale rewrites.</p><p>That balance — modernization without disruption — was the turning point. Swift Package Manager positioned us to evolve with the platform, rather than constantly working around it.</p><h3>Replacing a beating heart…</h3><p>This was open-heart surgery on a system that powered millions of app sessions every day — performed while the patient was still running around!</p><p>At the time of the migration, our ecosystem included:</p><ul><li><strong>~60 iOS/tvOS engineers </strong>actively shipping features</li><li><strong>~30 SDETs</strong> maintaining extensive automation and test infrastructure</li><li><strong>A multi-language codebase</strong> (Swift, Objective-C, and supporting tooling) spanning millions of lines of code</li><li><strong>60+ pods in total </strong>— external dependencies, internal SDKs, and local development pods under constant iteration</li><li><strong>Multiple markets</strong>(India, International) with distinct configurations</li></ul><p>Every one of those engineers relied on CocoaPods behaving in very specific, sometimes undocumented ways. Every automation script, every CI pipeline, every local development workflow had assumptions baked in about how dependencies resolved, how builds were structured, and how artefacts were produced.</p><p>We made extensive use of pre_install and post_install hooks. We maintained multiple build configurations for India and international markets — involving conditional linking, selective dependency inclusion, and market-specific code paths. Our CI pipelines were tightly coupled to CocoaPods-generated projects in ways that weren&#39;t always visible until something broke. This was organisational Jenga, only, we couldn’t let that tower fall, or even shake!</p><p>Easy, right?</p><h3>Core Constraint: Do Not Break Development</h3><p>From day one, we aligned on one non-negotiable rule:</p><blockquote><strong>This migration must not block ongoing development.</strong></blockquote><p>This constraint shaped everything.</p><p>Teams were actively shipping features. Local pods and internal SDKs were under constant iteration. CI pipelines had to remain stable across multiple environments. We couldn’t afford a <strong>“big bang”</strong> migration that froze development, forced widespread workflow changes, or introduced uncertainty into release cycles. Migration had to occur in phases with both systems co-existing.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/250/1*lfC9cdpEstTKcYtt6p9xMQ.gif" /></figure><p>This wasn’t the cleanest approach — maintaining two dependency systems in parallel added complexity. It was the safest path forward. It meant teams could keep shipping while we rebuilt the foundation underneath them.</p><h3>Our Migration Strategy</h3><p>Once we committed to a phased migration, we needed a structure that balanced safety with forward momentum. Each phase had to deliver real progress, while still preserving the ability to ship features without disruption.</p><p>We broke the journey into four deliberate stages:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qxtpM2tK4lFe-VnXDi05OA.png" /></figure><h4>Phase 1: Proving That SPM Could Coexist</h4><p>Before touching any critical paths, we focused on coexistence. Swift Package Manager was introduced alongside CocoaPods — not as a replacement, but as a parallel system. This phase was about validation.</p><h4>Phase 2: Internal Pods and Binary Distribution</h4><p>With coexistence proven, we moved to internal SDKs and binary dependencies. These were more controlled, lower-churn modules, making them ideal candidates for early migration. This phase helped us establish patterns for versioning, distribution, and consumption at scale.</p><p><em>This is where we built the playbook that would carry us through the harder phases ahead.</em></p><h4>Phase 3: Development Pods → Swift Packages</h4><p>This was the crucible.</p><p>Development pods were deeply intertwined with active feature work, often containing mixed-language code and UI resources. Migrating these required structural changes — not just mechanical conversions — and forced us to confront some of Swift Package Manager’s core constraints head-on.</p><p>This phase stretched across quarters, required constant coordination with feature teams, and tested every assumption we’d made about the migration strategy.</p><p><em>This is where patience mattered more than speed.</em></p><h4>Phase 4: Removing CocoaPods</h4><p>Only after the system had fully stabilized under SPM did we execute the final step: removing CocoaPods entirely. By this point, CocoaPods was no longer a dependency — it was technical debt waiting to be deleted.</p><p>The day we removed the Podfile from the repository felt like crossing a finish line we’d been racing toward for six months!</p><h3>The Payoff: Measurable Wins</h3><p>This wasn’t change for change’s sake. After two quarters of sustained effort, the results were tangible and significant.</p><h4>Immediate Impact</h4><ul><li><strong>Automation stability improved</strong> — test targets became less sensitive to implicit CocoaPods behavior</li><li><strong>Dependency graphs became explicit</strong> — making ownership, impact analysis, and refactoring significantly easier</li><li><strong>CI pipelines became more predictable</strong> — fewer pod-related cache invalidations and mysterious failures</li><li><strong>Build times improved</strong> <strong>by 2x </strong>— especially in incremental builds, due to better dependency isolation and binary targets wherever possible</li><li><strong>App startup time reduced by 200–300ms</strong> — SPM’s cleaner dependency loading eliminated redundant framework initialization at launch. The migration also allowed us to adopt the latest linker (previously blocked due to crashes on older OS versions), which improved dynamic library load times and static linking performance.</li></ul><h4>Long-Term Gains</h4><p>Just as importantly, we fundamentally reduced our risk profile. Dependency management stopped being a fragile layer propped up by scripts, conventions, and tribal knowledge. It became something the platform itself understood — native, supported, and evolving with the ecosystem.</p><p>We went from dreading the CocoaPods deprecation deadline to being ahead of it by a year.</p><p><strong>Six months. Two quarters. Sixty engineers. Sixty-plus pods. Zero feature freezes. One satisfying Podfile deletion.</strong></p><h3>The Scars — Part 2</h3><p>In <strong>Part 2</strong>, we’ll go deep into how these phases were executed in practice.</p><p>We’ll cover the real issues we encountered along the way — mixed Objective-C and Swift targets, XIBs in development pods, CI assumptions that quietly broke, and the architectural decisions we had to make to move forward safely.</p><p>This is where theory met reality — and where the migration truly earned its scars!</p><p>We’re hiring for our client teams! Do check out <a href="https://jobs.lever.co/jiostar?department=Digital+%7C+Engineering">open roles</a> if you want to build it right, while you build for millions of customers, who will use features that you build!</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=c3c7874566a9" width="1" height="1" alt=""><hr><p><a href="https://blog.hotstar.com/dependency-management-our-journey-beyond-cocoapods-part-1-the-strategy-c3c7874566a9">Dependency Management: Our Journey Beyond CocoaPods (Part 1 — The Strategy)</a> was originally published in <a href="https://blog.hotstar.com">JioHotstar</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
    </channel>
</rss>