Home Blog Page 397

Cambricon’s First Profit

The US-China AI Chip Race Enters a New Phase

Cambricon Technologies Reports First-Ever Quarterly Profit

The US-China AI chip race has entered a new phase as Chinese chip designer Cambricon Technologies reports its first-ever quarterly profit. The milestone emerges against a backdrop of escalating US export controls that have increasingly restricted Chinese companies’ access to advanced semiconductor technology, particularly Nvidia’s sophisticated AI processors.

A Shift in the US-China AI Chip Race

Cambricon’s breakthrough into profitability signals a significant shift in the US-China AI chip race, transforming from a 2016 startup into China’s most valuable artificial intelligence company, now valued at approximately 300 billion yuan ($41 billion). While this represents only a fraction of Nvidia’s $3 trillion market capitalization, it marks China’s growing capability to develop sophisticated AI chips domestically.

Financial Turnaround

The company’s financial turnaround is particularly noteworthy in the context of technological competition between the world’s two largest economies. After years of losses, Cambricon reported its first quarterly profit in the final quarter of 2024, with net profits ranging from 240 million yuan to 328 million yuan, despite posting a 724 million yuan loss in the first nine months.

Market Response

The market’s response to this shifting dynamic in the US-China AI chip race has been remarkable. Cambricon’s shares on the Shanghai Stock Exchange’s Star Market have surged more than 470% over the past year, climbing from 120.80 yuan to 695.96 yuan.

Growth Projections

The company projects a 70% revenue increase to 1.2 billion yuan in 2024, driven by China’s aggressive buildup of computing infrastructure to support its AI ambitions.

Technical Advancements

At the technical level, Cambricon has positioned itself as China’s answer to US chip restrictions with its 7-nanometre AI chips. The company’s flagship Cambricon-1A processor has gained significant traction in the domestic market, particularly in products from major technology companies like Huawei Technologies.

The Stakes in the US-China AI Chip Race

The stakes in the US-China AI chip race continue to rise, with analysts at Changjiang Securities projecting that China’s AI semiconductor market will reach 178 billion yuan by 2025. Beijing’s push for semiconductor self-sufficiency and increasing investments from domestic technology companies in AI infrastructure are fueling this growth.

Recent Developments

Recent US regulations announced in January 2025 have intensified the race, restricting Chinese access to advanced AI technology and limiting it to American companies and their allies. In response, major Chinese technology companies are investing heavily in domestic computing infrastructure.

Challenges and Opportunities

While Cambricon’s progress represents a significant advancement in the US-China AI chip race, challenges remain. The company must continue to narrow the technological gap with international competitors while maintaining its growth trajectory. However, supportive government policies and growing domestic demand provide a favourable environment for continued development.

Conclusion

Cambricon’s inclusion in the SSE 50 Index, which tracks the Shanghai Stock Exchange’s most valuable companies, underscores its strategic importance to China’s technology sector. As global tensions persist and access to foreign technology becomes more restricted, developing domestic AI chip capabilities has become increasingly important for China’s technological advancement and economic security.

FAQs

  • Q: What is the significance of Cambricon’s quarterly profit?
    A: Cambricon’s quarterly profit marks a significant shift in the US-China AI chip race, demonstrating China’s growing capability to develop sophisticated AI chips domestically.
  • Q: What are the implications of US export controls on the US-China AI chip race?
    A: The US export controls have restricted Chinese companies’ access to advanced semiconductor technology, particularly Nvidia’s sophisticated AI processors, intensifying the competition in the US-China AI chip race.
  • Q: What are the growth projections for China’s AI semiconductor market?
    A: Analysts at Changjiang Securities project that China’s AI semiconductor market will reach 178 billion yuan by 2025, driven by Beijing’s push for semiconductor self-sufficiency and increasing investments from domestic technology companies in AI infrastructure.

Microsoft Advances Materials Discovery with MatterGen

The Discovery of New Materials Just Got a Whole Lot Easier

The discovery of new materials is key to solving some of humanity’s biggest challenges. However, traditional methods of discovering new materials can feel like "finding a needle in a haystack."

Historical Methods of Materials Discovery

Historically, finding new materials relied on laborious and costly trial-and-error experiments. More recently, computational screening of vast materials databases helped to speed up the process, but it remained a time-intensive process.

Introducing MatterGen

Now, a powerful new generative AI tool from Microsoft could accelerate this process significantly. Dubbed MatterGen, the tool steps away from traditional screening methods and instead directly engineers novel materials based on design requirements, offering a potentially game-changing approach to materials discovery.

How MatterGen Works

Published in a paper in Nature, Microsoft describes MatterGen as a diffusion model that operates within the 3D geometry of materials. Where an image diffusion model might generate images from text prompts by tweaking pixel colors, MatterGen generates material structures by altering elements, positions, and periodic lattices in randomized structures. This bespoke architecture is designed specifically to handle the unique demands of materials science, such as periodicity and 3D arrangements.

A Leap Beyond Screening

Traditional computational methods involve screening enormous databases of potential materials to identify candidates with desired properties. Yet, even these methods are limited in their ability to explore the universe of unknown materials and require researchers to sift through millions of options before finding promising candidates.

In contrast, MatterGen starts from scratch—generating materials based on specific prompts about chemistry, mechanical attributes, electronic properties, magnetic behavior, or combinations of these constraints. The model was trained using over 608,000 stable materials compiled from the Materials Project and Alexandria databases.

Comparison with Traditional Screening Methods

In the comparison below, MatterGen significantly outperformed traditional screening methods in generating novel materials with specific properties—specifically a bulk modulus greater than 400 GPa, meaning they are hard to compress.

Experimental Synthesis of Novel Material

To prove MatterGen’s potential, Microsoft collaborated with researchers at Shenzhen Institutes of Advanced Technology (SIAT) – part of the Chinese Academy of Sciences – to experimentally synthesize a novel material designed by the AI.

Conclusion

Microsoft positions MatterGen as a complementary tool to its previous AI model, MatterSim, which accelerates simulations of material properties. Together, the tools could serve as a technological "flywheel," enhancing both the exploration of new materials and the simulation of their properties in iterative loops.

Frequently Asked Questions

Q: What is MatterGen?
A: MatterGen is a generative AI tool that directly engineers novel materials based on design requirements, offering a potentially game-changing approach to materials discovery.

Q: How does MatterGen work?
A: MatterGen operates within the 3D geometry of materials, generating material structures by altering elements, positions, and periodic lattices in randomized structures.

Q: How does MatterGen compare to traditional screening methods?
A: MatterGen significantly outperformed traditional screening methods in generating novel materials with specific properties.

Q: Has MatterGen been experimentally proven?
A: Yes, Microsoft collaborated with researchers at Shenzhen Institutes of Advanced Technology (SIAT) – part of the Chinese Academy of Sciences – to experimentally synthesize a novel material designed by the AI.

Biden Executive Order on Cybersecurity, AI, and More

0

US President Joe Biden Issues Sweeping Cybersecurity Directive

The 40-page executive order unveiled on Thursday is the Biden White House’s final attempt to kickstart efforts to harness the security benefits of AI, roll out digital identities for US citizens, and close gaps that have helped China, Russia, and other adversaries repeatedly penetrate US government systems.

Background

Looming over Biden’s directive is the question of whether president-elect Donald Trump will continue any of these initiatives after he takes the oath of office on Monday. None of the highly technical projects decreed in the order are partisan, but Trump’s advisers may prefer different approaches (or timetables) to solving the problems that the order identifies.

Key Provisions

The core of the executive order is an array of mandates for protecting government networks based on lessons learned from recent major incidents—namely, the security failures of federal contractors.

Software Vendors

The order requires software vendors to submit proof that they follow secure development practices, building on a mandate that debuted in 2022 in response to Biden’s first cyber executive order. The Cybersecurity and Infrastructure Security Agency would be tasked with double-checking these security attestations and working with vendors to fix any problems. To put some teeth behind the requirement, the White House’s Office of the National Cyber Director is “encouraged to refer attestations that fail validation to the Attorney General” for potential investigation and prosecution.

Cloud Platforms

The order gives the Department of Commerce eight months to assess the most commonly used cyber practices in the business community and issue guidance based on them. Shortly thereafter, those practices would become mandatory for companies seeking to do business with the government. The directive also kicks off updates to the National Institute of Standards and Technology’s secure software development guidance.

Internet-of-Things (IoT) Devices

To protect federal agencies from attacks that rely on flaws in internet-of-things gadgets, the order sets a January 4, 2027, deadline for agencies to purchase only consumer IoT devices that carry the newly launched US Cyber Trust Mark label.

Conclusion

The executive order is a significant step towards improving the government’s cybersecurity posture and ensuring the protection of its networks and data. While its effectiveness will depend on the implementation and enforcement of its provisions, it is a critical step towards securing the government’s digital foundations.

FAQs

Q: What is the purpose of the executive order?
A: The executive order aims to improve the government’s cybersecurity posture by requiring software vendors to follow secure development practices, assessing and mandating common cyber practices in the business community, and protecting federal agencies from attacks that rely on flaws in IoT devices.

Q: Will President-elect Trump continue these initiatives?
A: The order is not specific to a particular administration, and its provisions are designed to be continued or modified by future administrations.

Q: What is the US Cyber Trust Mark label?
A: The US Cyber Trust Mark label is a newly launched certification that indicates a consumer IoT device meets certain security standards and requirements.

Q: What is the deadline for agencies to purchase US Cyber Trust Mark labeled IoT devices?
A: The deadline is January 4, 2027.

Fantasy Gate Chronicles

0

and HTML: Understanding the Basics

What is

?

The

element is one of the most fundamental elements in HTML (HyperText Markup Language). It represents a paragraph of text and is used to define a block of text that forms a coherent thought or idea.

HTML Structure

The basic structure of the

element is as follows:

<p>Text goes here</p>

The opening

tag is used to indicate the start of the paragraph, and the closing

tag is used to indicate the end of the paragraph.

Usage and Examples

The

element is widely used in HTML documents to format text and separate it from other elements. Here are a few examples of its usage:

  • To define a paragraph of text:
    <p>This is a paragraph of text.</p>
  • To indent a paragraph of text:
    <p>   This paragraph is indented.</p>
  • To center a paragraph of text:
    <p style="text-align: center;">This paragraph is centered.</p>

    HTML5 Enhancements

In HTML5, the

element has undergone some changes and enhancements. Some of the key changes include:

  • The

    element is no longer required to have a closing tag. It can be self-closing, like this:

    <p>This is a self-closing paragraph.</</p>
  • The

    element can now contain other elements, such as headings, images, and links.

Conclusion

In conclusion, the

element is a fundamental element in HTML that is used to define a block of text. It is widely used in HTML documents to format text and separate it from other elements. Its usage and syntax have undergone some changes in HTML5, but its purpose remains the same.

Frequently Asked Questions

Q: What is the purpose of the

element?

A: The purpose of the

element is to define a block of text and separate it from other elements in an HTML document.

Q: Can the

element be self-closing?

A: Yes, in HTML5, the

element can be self-closing, but it is not required to be.

Q: Can the

element contain other elements?

A: Yes, in HTML5, the

element can contain other elements, such as headings, images, and links.

Silicon Valley defence start-up Shield AI hits $5bn valuation

Artificial Intelligence Start-up Shield AI to Nearly Double Valuation to $5bn

Raising $200mn from Defence and Aerospace Companies

Artificial intelligence start-up Shield AI will nearly double its valuation to $5bn in a new fundraising round, as investors rush to fund defence technology groups producing cutting-edge military systems.

Investors Line Up

The San Diego-based group, which makes AI-powered software for autonomous aircraft and drones, is raising about $200mn from defence and aerospace companies, including Palantir, Airbus, and L3 Harris, according to people close to the deal. Venture capitalists, including Andreessen Horowitz, Point72, and Riot Ventures, are also expected to participate.

Rationale for Investment

"Companies invest in competitors when there’s a strong strategic rationale to do so," said one of the people close to the matter. "The investments are a sign of the serious commitment to use our autonomy backbone in their programmes."

Growing Demand for AI in Defence

The fundraising comes as tech groups seek to grab a bigger slice of the US government’s $850bn defence budget from traditional prime contractors such as Lockheed Martin, Raytheon, and Boeing. Wars in Ukraine and the Middle East and geopolitical tensions between the US and China have heightened Washington’s reliance on tech companies developing advanced AI products for military purposes.

Shield’s Autonomous Technology

Shield’s core "Hivemind" software enables drones and aircraft to operate without GPS, communications, or a human pilot. Its autonomous technology is being used by a number of rival companies, including legacy defence prime contractors that can incorporate its software into their aircraft.

Conclusion

The surge in investment in Shield AI and other defence technology companies reflects the growing importance of AI in military applications. As the US government continues to increase its spending on national security, it is likely that we will see even more investment in this space in the coming years.

FAQs

Q: What is Shield AI’s core technology?
A: Shield’s core "Hivemind" software enables drones and aircraft to operate without GPS, communications, or a human pilot.

Q: Who is investing in Shield AI?
A: Palantir, Airbus, L3 Harris, Andreessen Horowitz, Point72, and Riot Ventures are among the companies investing in Shield AI.

Q: Why are investors interested in Shield AI?
A: Investors are interested in Shield AI because of its potential to provide cutting-edge military systems and its ability to integrate with other defence companies.

ZDNET’s Pick for Best Robot Vacuum and Mop is Nearly Half Off

What’s the Deal?

Our pick for the best robot vacuum and mop combination, the Dreame X40 Ultra, is now discounted by $900 off its original price, on sale for $1,000.

ZDNET’s Key Takeaways

* The Dreame X40 Ultra is typically $1,900, but it’s 42% off right now on Amazon, available for $1,000.
* The X40 Ultra is a high-performing robot with excellent mapping capabilities and strong 12,000Pa suction. It performs exceedingly well on carpet and hard floors, with great object avoidance and high customization.
* Although it can recognize and show snapshots of obstacles, it sometimes gets tangled in cords and loses one or both mop pads. I also found some connectivity issues with the app.

As a fan of robot vacuums and mops, I’m always looking for the next big thing in home cleaning robotics. I almost stopped looking after the Dreame X40 Ultra.

Setup was quite a breeze — you download the Dreamehome app and follow the instructions to add it to your Wi-Fi network and set up your preferences. I found some issues with the app connecting to the Dreame X40 Ultra, making me wait a few minutes or force-quit the Dreamehome app to relaunch it. Even a month later, I found that this still happens occasionally, which is disappointing as I don’t have that issue with other robots.

The robot features market-leading suction power at 12,000Pa, making it excellent for carpeted homes or homes with pets. It’s efficient and fast, easily cleaning across my floors and navigating obstacles while staying true to its map.

One of my two favorite features of the Dreame X40 Ultra robot is its object avoidance. This is one of two robots in my home that doesn’t require me to pick up every object from the floor before running it. This means I can send it out to clean on a schedule or when I’m away from home and rely on it to consistently deliver a clean home without getting stuck on a kid’s sock under the coffee table.

My second favorite thing about this robot is that it has magnetic mop pads that not only lift to avoid getting carpets wet but can also be set up in the app to leave the mop pads at the base, vacuum the carpets first, and then vacuum everything else. Then, it returns to the base to reattach its mop pads and mop the entire house, avoiding the already clean carpets.

Unfortunately, the X40 Ultra’s mop pads lift to about 10.9 mm, which is not quite enough to keep my living room’s medium pile carpet dry. This means I have to set that rug as a no-go zone for robot vacuum and mop combinations like the Yeedi M12 PRO+, which lifts its mop pads to 9 mm.

The Dreamehome app has so many customizations that it reminds me of everything you can do with the Roborock S8 MaxV Ultra, a direct high-end competitor with many of the same features. Like the S8 MaxV Ultra, the X40 Ultra takes photos of obstacles and can take photos of your pets in passing if it spots them.

Dreame uses AI for visual recognition to identify and add objects to your map. Of course, this isn’t always accurate, which is why my toddler was mistaken for a pet, but it’s pretty entertaining and useful that the robot lets you see the obstacles it finds in its path. Using your smartphone, you can also drop into the camera’s feed to view what the robot sees,

The X40 Ultra’s camera can detect stains on hard floors and rev up the robot’s mopping power to scrub up the stains. When this happens, the side brush automatically lifts to avoid getting wet or spreading wet messes. When the built-in turbidity sensor detects too much dirty water, the robot returns to the base station to rewash its mop pads and resumes its cleaning session.

How is this robot not perfect? To address the elephant in the room, the Dreame X40 Ultra is one of the most expensive robot vacuum and mop combinations I’ve seen, at $1,900. Its current limited-time deal helps a lot. The Roborock S8 MaxV Ultra is priced at $1,800, and I already find that to be steep. Spending almost two grand on a robot that will roll around your dirty floors isn’t an easy purchase.

It’s also simply not perfect because nothing is. I found that the X40 Ultra often tries to go over some extension cords rather than avoid them, which almost always results in one or both of its mop pads coming off and the robot getting stuck without them. This is a bigger inconvenience when I’ve left it alone to clean the house and come back to see it surrounded by dirty floors in a corner.

ZDNET’s Buying Advice

When it comes down to it, the Dreame X40 Ultra is the smartest robot vacuum and mop I’ve ever tested.

One simple example: If you’ve ever had robot vacuums, you’ve likely experienced them aimlessly roaming around when it’s time to return to the dock, only to pause two feet away to say they can’t find the charging station.

This Dreame X40 Ultra never does that. As a robot vacuum reviewer, I decided to bring out 11 robot vacuums in a group to take photos of them. After the photos, the X40 Ultra accidentally got bumped when someone tried to walk over the robot labyrinth and began returning to the dock. After expertly navigating through 10 of its pals, weaving to and fro like a cyclist through a traffic jam, it went straight to the charging dock and began charging. I tried this with three other robots, but not one found its charging dock.

I have three little kids set on making as many messes as possible, and the fact that I don’t have to worry about this robot getting its roller brush stuck makes the Dreame X40 Ultra one of the best robot vacuum and mops I’ve ever used.

Conclusion

ZDNET’s team of experts constantly monitors the deals we feature to keep our stories up-to-date. If you missed out on this deal, don’t worry – we’re always sourcing new savings opportunities at ZDNET.com.

FAQs

Q: What is the Dreame X40 Ultra?
A: The Dreame X40 Ultra is a robot vacuum and mop combination that combines advanced navigation and mapping capabilities with strong suction power and customizable cleaning settings.

Q: How much does the Dreame X40 Ultra cost?
A: The Dreame X40 Ultra is typically $1,900, but it’s currently on sale for $1,000.

Q: What are the key features of the Dreame X40 Ultra?
A: The Dreame X40 Ultra features market-leading 12,000Pa suction power, object avoidance, and customizable cleaning settings, as well as a built-in turbidity sensor and camera for detecting and cleaning stains.

Q: How does the Dreame X40 Ultra compare to other robot vacuum and mop combinations?
A: The Dreame X40 Ultra is one of the most advanced and feature-rich robot vacuum and mop combinations on the market, with a strong focus on navigation, mapping, and customization. It’s a top choice for those looking for a high-performing and reliable cleaning solution.

Donkey Kong’s Adorable New Design

0

Nintendo Switch 2: Donkey Kong’s Revamped Look Sparks Debate

A New Direction for Donkey Kong’s Design

By now you’ve likely seen the highly anticipated Switch 2 launch trailer featuring all-new Mario Kart gameplay, but some eagle-eyed viewers have noticed a strange detail with Donkey Kong. Cuter and goofier than ever, the new design is a far cry from what we’re used to (and surprisingly, I don’t hate it).

A Blast from the Past: The OG Donkey Kong Design

Since his debut, Donkey Kong’s character design has remained fairly unchanged, but his ‘new’ look isn’t quite as fresh as you might expect. Seemingly taking inspiration from the animated aesthetics of the 2023 Super Mario Bros. Movie, DK’s revamped look is dividing classic Nintendo fans.

A Shift towards Cuteness

The Rare-era Donkey Kong design featured in classics like Donkey Kong County will be sorely missed by fans, but (and don’t hate me when I say this) I find the ‘new’ design quite endearing. With a less angular brow and softer appearance, the Switch 2 DK is reminiscent of the OG arcade design. Sure, he’s a little goofy, but the exaggerated proportions and animated appeal bring a new dimension and welcome dose of personality to his character.

The Cuter Look of Mario Kart Characters

While DK’s revamp is the most noticeable, all the Mario Kart characters have a cuter look compared to the older games, with some fans noting the tweaks to Mario’s design resemble his Super Mario Bros. Wonder look. "They all look way cuter than they did in Mario Kart 8, love this style," one X commenter wrote. Another added, "I really like just how cartoony the art style looks for Mario Kart 9… look at Donkey Kong, he looks so silly."

A Shift in Tone

RIP Rare Donkey Kong 1994-2025 You’ve lived a nice good life buddy. pic.twitter.com/4Q7J0hPQAF

Now I know why Nintendo tweeted so much about Donkey Kong Country HD It was to mentally prepare us to see this for the first time pic.twitter.com/KcNKIXoskt

Conclusion

Donkey Kong’s revamped look may not be to everyone’s taste, but it’s undeniable that it brings a new level of personality and charm to the character. As fans continue to debate the new design, it’s clear that Nintendo is pushing the boundaries of their iconic characters.

Frequently Asked Questions

Q: What is the inspiration behind Donkey Kong’s new design?
A: The new design appears to be inspired by the animated aesthetics of the 2023 Super Mario Bros. Movie.

Q: Will the Rare-era Donkey Kong design make a comeback?
A: Unfortunately, it seems unlikely, as Nintendo has opted for a new direction with Donkey Kong’s character design.

Q: Will other Mario Kart characters receive similar design changes?
A: Yes, all the Mario Kart characters have received a cuter look compared to the older games.

Revolutionizing Educator Growth with Generative AI

The Role of Generative AI in Educator Professional Development

The Need for AI in Professional Development

Educators often struggle to balance professional training with the constraints of time, accessibility, and relevance. Professional development programs often rely on a one-size-fits-all model, which cannot by definition address the needs of each individual teacher. Generative AI can assist in enabling a new era for professional educator growth.

AI in Personalized Professional Growth

One of the most promising applications of generative AI is its ability to design individualized professional development pathways. Acting as a training planner, AI tools can analyze an educator’s previous experiences, current skill set, teaching context, and long-term career goals to recommend targeted learning opportunities at all stages of the educator’s journey.

Streamlining Continuing Education Tracking

Educators are often required to earn continuing education units (CEUs) to maintain their licenses, an often-complicated process governed by unfamiliar bureaucratic requirements. Generative AI can streamline this process by automating the tracking and documentation of CEUs. Imagine an AI-integrated system that logs professional development activities as they occur, calculating and updating CEU credits in real-time.

Challenges and Ethical Considerations

While the benefits of generative AI in PD ecosystems are exciting, it is important to approach these tools with a critical eye. Potential challenges include data privacy concerns, the risk of over-reliance on AI recommendations, and the need to address bias within AI algorithms. Gullani, et al. recently wrote that schools and districts must ensure that AI-enhanced tools are transparent, equitable, and aligned with their educational mission. Educators must always view AI as a supplement to, not a replacement for, human expertise.

Conclusion

Generative AI is redefining educator professional development by offering tools and strategies to create smarter, more impactful, and targeted learning opportunities. From personalized PD plans to AI-supported assessment and feedback mechanisms, these innovations streamline processes, save educators valuable time, and provide actionable insights to support growth.

FAQs

Q: How does AI support personalized professional development?
A: AI can design individualized professional development pathways, recommending targeted learning opportunities based on an educator’s previous experiences, current skill set, and long-term career goals.

Q: How does AI streamline continuing education tracking?
A: AI can automate the tracking and documentation of CEUs, logging professional development activities as they occur and calculating and updating CEU credits in real-time.

Q: What are the challenges and ethical considerations of using AI in PD?
A: Potential challenges include data privacy concerns, the risk of over-reliance on AI recommendations, and the need to address bias within AI algorithms. Educators must always view AI as a supplement to, not a replacement for, human expertise.

Introducing KV Cache Reuse Optimizations in NVIDIA TensorRT-LLM

0

Priority-based KV Cache Eviction

When an LLM request has completed, the KV cache blocks associated with these requests are stored. Given the bounded size of the KV cache, some cached blocks may need to be evicted to make room for new sequences. By default, eviction follows a least recently used (LRU) policy.

Priority-based eviction is a new feature of the TensorRT-LLM Executor API that enables users to influence how blocks are selected for eviction. Users can specify two attributes that guide block eviction: priority and duration. The priority value sets the relative retention priority (how important it is to retain that block in the cache), and the duration value sets how long this priority level should apply for.

The priority based-eviction API enables an LLM deployer to use knowledge about their workload to improve reuse opportunities by persisting blocks that are likely to be reused. For example, the deployer may want blocks corresponding to a system prompt to stay in the cache as long as possible, or blocks that might be involved in a latency-critical request should persist with higher priority than others (Figure 1).

For each request, you can specify a priority and duration value for discrete ranges of tokens in the input context, along with a priority and duration for blocks allocated during the decode phase. The priority level of a range of tokens applies until the duration has passed after no period of reuse, or until the blocks corresponding to these ranges have been evicted.

When choosing blocks to be evicted, TensorRT-LLM considers the priority levels of tokens within the block. For example, a request with a 500-token system prompt can set the token range [0, 500) to the maximum priority. This way, the cache blocks corresponding to these tokens will only be evicted if absolutely necessary. Alternatively, if you know that blocks will never be reused, you can set the blocks of this request to the lowest priority to ensure that they are evicted first, before other blocks.

This new implementation also biases toward blocks further from the root, which leads to a small performance improvement, even when not setting priority levels. Our internal benchmarks show priority-based eviction increasing cache hit rate by around 20% and varies based on the workload.

# Priority-based eviction usage examples

#Example 1: One-off request

KvCacheRetentionConfig(
[TokenRangeRetentionConfig(start=0, end=null, priority=0)],
decode_priority=0
)

#Example 2: High Priority system prompt

KvCacheRetentionConfig(
[TokenRangeRetentionConfig(start=0, end=1000, priority=100)]
)

#Example 3: Retain context blocks for 30 seconds, and decode blocks for 10 seconds

KvCacheRetentionConfig(
[TokenRangeRetentionConfig(start=0, end=null, priority=100, duration=30s)],
decode_priority=100, decode_duration=10s)

KV Cache Event API

In large-scale LLM-powered applications, deployers often provision multiple serving instances of a model to distribute incoming requests. This raises the question, which instance should process new requests? Requests are often routed to balance load to ensure efficient utilization and quick processing of any request. The size of the KV cache on any instance represents the capacity to grow and accept new work.

However, load-based routing may not be optimal. If a moderately loaded instance has already computed and cached the keys and values for a new request, routing the request to this instance might still be preferred to optimize for cache reuse. The KV cache event API enables request routing systems to track which instances have cached or evicted blocks, enabling more intelligent reuse and greater performance.

The TensorRT-LLM Executor API now exposes a means of tracking updates to the KV cache.

# Set the max size of the internal event buffer. Defaults to 0 (no events)
kv_cache_config = KvCacheConfig(event_buffer_max_size=16384)

executor_config = ExecutorConfig(kv_cache_config)

executor = Executor(executor_config)

# Get an event manager
eventManager = executor.getKvCacheEventManager()

# Wait for new events. Once it returns, it implicitly clears the internal queue of events. Optionally provide a timeout value. If there’s no events within this timeout, it returns an empty list.
events = eventManager.getLatestEvents()

When a cache block is stored for reuse, removed, or updated, an event is emitted. These events can be consumed in real time by an application to get an eventually consistent view of the current state of the TensorRT-LLM KV cache. This is especially useful for tracking KV cache reuse opportunities. It can be used on the scale of a single executor to anticipate which requests will have more reuse, or aggregated across many executors to make KV-aware routing and scheduling decisions (Figure 2).

# KV cache event API scalable implementation where events are processed across KV tree shards and aggregated across multiple executors to provide KV-aware routing and scheduling of requests that optimize KV cache reuse opportunity.

Summary

NVIDIA TensorRT-LLM provides several optimizations to efficiently deploy your generative AI applications across NVIDIA-accelerated infrastructure anywhere, including cloud, data center, and workstations. These optimizations lead to significant speedups and better cache reuse on the same hardware. This ultimately enables using fewer resources to serve the same workload, reducing energy costs, and improving total cost of ownership.

Nexos.ai Helps Enterprises Scale AI Projects from Pilot to Production

0

Capitalizing on a Catalyst

A new AI orchestration startup from the founders of Lithuanian unicorn Nord Security is setting out to help enterprises put their AI projects into production, with an initial focus on bringing greater visibility, security, and adaptability to large language models (LLMs).

Capitalizing on a Catalyst

Currently, teams that want to put their AI into production have to connect myriad tools, which likely involves recruiting and building teams with the necessary skills. This is where Nexos.ai wants to step in.

The Idea Behind Nexos.ai

Nexos.ai’s founders, Tomas Okmanas and Eimantas Sabaliauskas, have a deep understanding of the challenges that come with deploying AI models in production. They built one of the most recognizable brands not only in Lithuania but in all of Europe, Nord Security, best known for its flagship VPN product NordVPN.

The Problem with AI in Production

Okmanas has seen firsthand the difficulties that come with putting AI into production. “I’ve seen that there’s a big gap between running AI as pilots and going into production,” he said. “When you’re testing AI in your lab, it might work and it can be useful, but when you want to put it into production, especially in enterprises, how do you ensure high availability? How do you ensure security? How do you manage cost?”

The Solution: Nexos.ai

Nexos.ai is a platform that provides a simple API (application programming interface) for accessing more than 200 AI models, from big-name incumbents like OpenAI and Anthropic to smaller, niche LLMs. The idea is that if OpenAI goes down, a company can temporarily (and automatically) switch to a different provider without breaking stride.

Security and Compliance

Nexos.ai also ushers “intelligent caching” into the mix — if a particular question is repeated by multiple users, the system can turn to its own database rather than continuing to engage the LLM, which can get expensive. On the security and compliance fronts, Nexos.ai prevents individuals from sending private data to LLM providers, or if an employee leaves a company, their access can be terminated immediately.

From Idea to Inception

Going from an idea to formal incorporation took Nexos.ai around six weeks, and while the speed of securing the funding was largely down to the founders’ pedigree, a big part of it was simply the timing.

The Future of Nexos.ai

Nexos.ai’s platform is set to launch by the end of March, though Okmanas said it is already working with a bunch of “beta customers and design partners.” The company will likely offer self-hosting in the future, and it already supports integrations with companies’ own internal LLMs.

Conclusion

Nexos.ai is poised to revolutionize the way enterprises deploy and manage AI models in production. With its simple API, intelligent caching, and focus on security and compliance, the platform is well-positioned to help companies overcome the challenges of putting AI into production.

Frequently Asked Questions

Q: What is Nexos.ai?

A: Nexos.ai is a platform that provides a simple API for accessing more than 200 AI models, from big-name incumbents like OpenAI and Anthropic to smaller, niche LLMs.

Q: What is the problem with AI in production?

A: The problem with AI in production is that it can be difficult to ensure high availability, security, and cost management when deploying AI models in production.

Q: How does Nexos.ai solve this problem?

A: Nexos.ai solves this problem by providing a simple API for accessing AI models, intelligent caching, and a focus on security and compliance.

Q: When is Nexos.ai launching?

A: Nexos.ai is set to launch by the end of March, though the company is already working with beta customers and design partners.