Popular Posts

LinkedIn Bucks AI Infrastructure Boom, Prioritizing Efficiency Over Aggressive Expansion

Unlike many of its big tech counterparts, LinkedIn, the professional social network with over 1.3 billion users, has made a strategic decision to significantly rein in its spending on artificial intelligence infrastructure this fiscal year. Instead of aggressively expanding its AI data centers, the company plans to maintain a steady investment in Graphics Processing Units (GPUs) and keep its overall compute and storage footprint flat. This move stands in stark contrast to the prevailing trend among major tech players who are pouring vast sums into building out their AI capabilities.

The spending calculations for LinkedIn apply to its fiscal year, which commenced last month and is set to conclude in June of next year. Executives at the Microsoft-owned platform revealed to WIRED that this ambitious strategy is possible because the company has successfully boosted the efficiency of its existing GPUs by twofold over the past six months. While acknowledging the rapidly shifting hardware demands of AI could potentially challenge this plan, LinkedIn’s leadership asserts that the company has already factored in the surging prices of critical components like memory chips into its projections.

Erran Berger, LinkedIn’s chief technology officer for engineering, articulated the company’s bold objective: "One of the goals we’ve set is to try to basically keep our compute footprint flat or as close to flat as possible while shipping more compute-hungry things to production." He added that making such a statement in the current technology landscape is "pretty bold." Berger, alongside Raghu Hiremagalur, LinkedIn’s chief technology officer for infrastructure, emphasized a commitment to prudent spending. They believe these new constraints will foster greater creativity among engineering teams as they develop and roll out the numerous generative AI features LinkedIn intends to launch. Berger also expressed confidence that these efficiency gains could compound over time, allowing LinkedIn to extract even more value from future data center expansions when budgets eventually increase.

Hiremagalur underscored the magnitude of this undertaking for a company of LinkedIn’s scale. "I really want to double underscore that for a company of our scale, to say a full year we’re going to do this with no incremental storage and compute is no small feat, but it’s taken a ton of work to get there," he stated, highlighting the intensive effort required to achieve such a goal.

This approach sharply diverges from the strategies of other industry behemoths such as OpenAI, Meta, and Google, which are actively consolidating financial resources and forging unexpected partnerships to construct, equip, and operate massive new data centers. These facilities are designed to house the latest and most powerful computer chips, essential for training and running advanced AI models. The intense demand has led to widespread labor and parts shortages, causing delays in many projects and even forcing some businesses to limit customer access to certain AI tools. Amidst this frenetic expansion, there are growing questions about the long-term sustainability—both economic and environmental—of such relentless investment in AI infrastructure. LinkedIn, by bucking this building boom, becomes perhaps the largest business yet to publicly address these spending concerns through a deliberate focus on optimization rather than sheer scale.

Songyee Yoon, managing partner of Principal Venture Partners and a board member at server manufacturer HP, views LinkedIn’s decision as a positive indicator for the broader industry. "It is encouraging for the industry," Yoon commented. "It suggests AI is beginning to move from experimentation into production discipline. The companies that win will not simply be the ones that spend the most on infrastructure." This perspective suggests a maturing of the AI landscape, where strategic resource management may prove as crucial as raw computational power.

LinkedIn’s journey to this point is rooted in its infrastructure history. A few years after Microsoft acquired the company in 2016, LinkedIn explored migrating its operations to its parent company’s Azure cloud service. However, this proved to be economically unfeasible. As Hiremagalur explained, "Microsoft Azure was growing like crazy, the level of customer demand was through the roof, and at the same time we saw skyrocketing growth on the LinkedIn side." Attempting to squeeze the colossal social network into general-purpose data centers within Azure did not make financial sense given both companies’ rapid expansion.

Consequently, in 2022, LinkedIn made the strategic decision to go "all-in" on its own dedicated data centers, establishing and operating facilities in key locations such as Oregon, Texas, and Virginia. This move granted LinkedIn significant control over every facet of its technology stack, positioning it well to adapt to the evolving demands of a new era dominated by AI. Around the same time, LinkedIn began the ambitious endeavor of developing AI-based assistants designed to enhance user experience by helping with tasks like message composition, job searching, and candidate recruitment. This undertaking, however, came with escalating costs. Hiremagalur noted the unsustainable trajectory, stating, "Every query that’s coming to our site has increased in cost over time," and highlighted that the sheer volume of data LinkedIn stored was doubling annually. "That is not a sustainable place to be," he concluded.

Driven by these escalating costs and the imperative to manage its growing data footprint, LinkedIn embarked on a comprehensive initiative to optimize its data center usage across the entire AI pipeline, from the initial training of complex models to serving them efficiently in response to user queries. Hiremagalur’s team developed sophisticated measurement tools to meticulously track the exact amount of compute and storage resources consumed by individual engineering teams. This granular understanding enabled them to establish a system for allocating projects to the various computers within its data centers with unprecedented efficiency, significantly reducing idle time. Hiremagalur proudly described their success, stating, "Our allocation efficiency and utilization of GPUs on the training side is the best that I have seen," reporting usage rates "at north of 95 percent."

Beyond efficient allocation, LinkedIn also employed advanced techniques such as model distillation. This process involves training smaller, more resource-efficient AI models by leveraging the knowledge and insights gleaned from larger, more complex ones. For instance, in its job recommendation tools, LinkedIn developed a single, smaller model that learned from two larger models. This streamlined model became proficient in both identifying relevant job openings and accurately predicting which users were most likely to click on them. Berger affirmed that despite its smaller size and lower operational cost, this distilled model does not compromise on quality. "People are finding and discovering jobs that they were not successfully finding before, because the model is doing a really good job of understanding" their preferences and desires, he explained.

Similarly, the AI model responsible for selecting which posts appear in users’ newsfeeds, initially very expensive to operate, was also optimized to run in "a reasonably cost-effective way," according to Berger. This was achieved through dozens of improvements, including streamlining model training processes, reusing information from earlier recommendations to avoid redundant computations, and expertly balancing workloads between less expensive Central Processing Units (CPUs) and the more costly, power-intensive GPUs.

LinkedIn even took the extraordinary step of reworking some of the foundational software that runs on Nvidia processors, including open-source projects like the Liger kernel. This allowed them to make these processors capable of handling tasks larger than their original design specifications. Additionally, they rejiggered other software components to offload tasks from Nvidia GPUs—which are notoriously pricey, difficult to procure, and consume substantial amounts of electricity—and run them on more readily available and energy-efficient CPUs. Cumulatively, LinkedIn estimates that these extensive efficiency efforts have resulted in approximately $24 million in savings over the past 12 months, a figure equivalent to the cost of running roughly 1,100 GPUs continuously for a full year.

While acknowledging that $24 million might not seem like a colossal sum for a company boasting $18 billion in annual sales, Hiremagalur emphasized that "craft" and "agility" are equally important. By freeing up computing resources, engineers gain the flexibility and capacity to embark on their next projects sooner and integrate more sophisticated AI capabilities without the immediate need to expand LinkedIn’s computing footprint. Berger further contended that these efforts enable LinkedIn to deliver superior job and candidate matching results, deploying larger models and performing deeper inference, all while keeping costs contained and generating a healthy financial return for Microsoft. "We should be able to deliver better quality by deploying larger models, doing deeper inference, and doing it for cheaper if we can," he asserted.

Despite the flat spending on expansion, LinkedIn’s data centers are far from becoming obsolete. The company has proactively committed to purchasing new servers, ensuring it can continually upgrade its existing machines as they age out or encounter breakdowns in the coming months. This forward-thinking procurement strategy also allowed LinkedIn to lock in significant savings by buying hardware ahead of further price increases. Hiremagalur lamented the current market volatility, stating, "The cost of all of this hardware has just gone through the roof," noting that prices for some servers have inexplicably jumped threefold in just a few months. "It’s just nuts," he remarked.

LinkedIn’s efficiency drive aligns with a broader industry trend sometimes referred to as "tokenomics," which involves a more in-depth analysis and optimization of the costs associated with using generative AI tools. Chirag Dekate, a consultant at Gartner who advises businesses on their AI cloud strategies, observed, "Enterprises are evolving from a buy-more era to a do-more era. Until now, the mantra was, buy more to save more. But buy more only increases costs." Dekate noted that even smaller businesses, which typically lack the extensive control over their infrastructure that LinkedIn possesses, are finding innovative ways to reduce costs. These methods include purging unused software, streamlining employee headcount, purchasing data center space from "neoclouds"—providers often more affordable than traditional cloud giants—and adopting the most cost-effective AI models available for specific projects.

However, Dekate also voiced a cautionary note, expressing concern that LinkedIn and other companies implementing stringent financial constraints might eventually encounter a "wall." He believes that the fundamental need for ever-greater amounts of compute and storage in the evolving AI landscape is ultimately inevitable. "At some point, something has to give," Dekate warned, implying that companies might eventually have to compromise either on their AI ambitions or their mandates to freeze IT spending.

LinkedIn is not entirely ruling out future data center growth. However, it is fundamentally altering its approach to infrastructure allocation. The era of making "a wild ass guess" about the computing resources teams might need for the coming year is over, according to Hiremagalur. Instead, LinkedIn has chosen to "embrace the chaos" of the rapidly changing AI landscape, taking things quarter by quarter while maintaining a steadfast focus on long-term return on investment, Berger explained. While the current flattened spending plan is ambitious, the executives acknowledge that it might not last precisely as long as initially projected, reflecting the dynamic nature of AI development and infrastructure needs.

Leave a Reply

Your email address will not be published. Required fields are marked *