Home Blog Page 520

Digital Pharmacy Ordering Enabled in Victoria

NSW Extends Virtual Health Service to Adults

Over a year since offering urgent virtual care service to children and youth statewide, the New South Wales government has come up with a similar service for adults.

New VirtualAdult Service

The new VirtualAdult service will provide urgent care via video conferencing for common illnesses or injuries, such as coughs and colds or flu, respiratory symptoms, vomiting and diarrhoea, and rashes.

Statewide Expansion

Following its launch this month, December, in Sydney, it is planned for statewide expansion by the end of 2025. The service is part of several initiatives introduced to relieve pressure on overwhelmed emergency departments across NSW.

Digital Pharmacy Ordering Now Live in Victoria

Compounding pharmacy Slade Health, part of cancer care provider Icon Group, has digitised its ordering system for public hospitals in Victoria.

Automation of Information Flow

It recently integrated the digital medications management system, Charm Evolution by Magentus, with its in-house software, to automate information flow.

Benefits of the Digital System

In a media release, Slade Health says public health facilities and hospitals can now order through this digital system, doing away with manual entry. The system also provides the pharmacy with near real-time visibility into orders, updates, and cancellations.

Full Digitization

This integration makes Victoria the second state in Australia after Queensland to fully digitise ordering for hospital pharmacies.

Royal Melbourne Hospital Beefing Up Digital Transformation

Royal Melbourne Hospital has sought support from digital health accelerator ANDHealth for its digital transformation.

Innovation Partner

It made the organisation its “innovation partner,” helping keep updated and providing expert advice on new and emerging technologies.

Core Hospital System Partner

The hospital has also become ANDHealth’s first core hospital system partner, providing insights into clinician and staff demand and needs from technology.

Conclusion

The NSW government’s extension of virtual health services to adults, the digitization of pharmacy ordering in Victoria, and Royal Melbourne Hospital’s digital transformation efforts demonstrate the growing importance of digital healthcare in Australia.

FAQs

Q: What is the VirtualAdult service?

A: The VirtualAdult service is a new urgent care service via video conferencing for common illnesses or injuries, such as coughs and colds or flu, respiratory symptoms, vomiting and diarrhoea, and rashes.

Q: When will the VirtualAdult service be available statewide?

A: The service is planned for statewide expansion by the end of 2025.

Q: What is the digital medications management system used by Slade Health?

A: The digital medications management system used by Slade Health is Charm Evolution by Magentus.

Q: What are the benefits of the digital system used by Slade Health?

A: The digital system provides public health facilities and hospitals with the ability to order through a digital system, doing away with manual entry, and provides the pharmacy with near real-time visibility into orders, updates, and cancellations.

SuperNICs for Next-Generation AI Networking

0

Leveraging RoCE for AI Workloads

For AI model training, it is critical to move immense datasets at high speed between GPUs across the data center to reduce training time and achieve faster time-to-market for AI solutions.

NVIDIA SuperNICs, featuring best-in-class, in-hardware RoCE acceleration and GPUDirect RDMA at speeds up to 800 Gb/s, address these challenges by enabling direct data movement between GPUs while bypassing the CPU.

Enhancing AI Performance with Spectrum-X RoCE Adaptive Routing

One of the key capabilities for boosting AI network performance within Spectrum-X is direct data placement (DDP) support featured by NVIDIA SuperNICs.

Spectrum-X RoCE adaptive routing dynamically adjusts how traffic is distributed across available network paths, ensuring that high-bandwidth flows are optimally routed to prevent network congestion.

Addressing Congestion in AI Networks

AI workloads are highly susceptible to congestion due to their bursty nature. The frequent, short-lived traffic spikes generated by AI model training—particularly during collective operations where multiple GPUs synchronize and share data—require advanced congestion management to maintain network performance.

To address this, Spectrum-X employs advanced congestion control mechanisms that are tightly integrated with the Spectrum-4 switch’s real-time telemetry capabilities.

Accelerating AI Networks with Enhanced Programmable I/O

As AI workloads grow more complex, network infrastructure must evolve not only in speed but also in adaptability to support diverse communication patterns across thousands of nodes.

NVIDIA SuperNICs are at the forefront of this innovation, offering enhanced programmable I/O capabilities that are crucial for modern AI data center environments.

Securing Network Connectivity for AI

Securing AI models is essential for protecting sensitive data and intellectual property from potential breaches and adversarial attacks.

Traditional network encryption methods often struggle to scale beyond 100 Gb/s, leaving critical data at risk. In contrast, NVIDIA SuperNICs offer accelerated networking with in-line crypto acceleration at speeds of up to 800 Gb/s, ensuring that data remains encrypted in transit while achieving peak AI performance.

Conclusion

In the dynamic landscape of generative AI, NVIDIA SuperNICs are setting the stage for a transformative era in networking, serving as an integral part of the NVIDIA Spectrum-X and Quantum-X800 networking platforms.

With their unparalleled capabilities—from ultra-fast data throughput and intelligent congestion management to robust security features and programmable I/O—these network accelerators are revolutionizing how AI workloads are delivered.

FAQs

Q: What is the purpose of NVIDIA SuperNICs?

A: NVIDIA SuperNICs are designed to power hyperscale AI workloads by providing state-of-the-art Ethernet and InfiniBand solutions that maximize the performance and efficiency of AI factories and cloud data centers.

Q: What is RoCE acceleration, and how does it benefit AI workloads?

A: RoCE acceleration enables direct data movement between GPUs while bypassing the CPU, reducing latency and improving overall system efficiency.

Q: How does Spectrum-X RoCE adaptive routing improve AI network performance?

A: Spectrum-X RoCE adaptive routing dynamically adjusts how traffic is distributed across available network paths, ensuring that high-bandwidth flows are optimally routed to prevent network congestion.

Q: How do NVIDIA SuperNICs address congestion in AI networks?

A: NVIDIA SuperNICs employ advanced congestion control mechanisms that are tightly integrated with the Spectrum-4 switch’s real-time telemetry capabilities, enabling proactive adjustments to data transmission rates based on current network utilization.

Q: What is the significance of programmable I/O in AI networks?

A: Programmable I/O enables network professionals to build and optimize networks at massive scale, providing the flexibility to tailor network infrastructure to the specific needs of AI workloads.

Q: How do NVIDIA SuperNICs secure network connectivity for AI?

A: NVIDIA SuperNICs offer accelerated networking with in-line crypto acceleration at speeds of up to 800 Gb/s, ensuring that data remains encrypted in transit while achieving peak AI performance.

Quick Gestures

Gemini: The Google Assistant Replacement Gets a Major Update

Gemini has been available as a Google Assistant replacement for a while now. At first, Gemini wasn’t an especially helpful assistant because it lacked many required features, but Google has been slowly rolling out new features to help round out Gemini.

New Feature Rollout

As reported on 9to5Google, the most recent feature rollout includes the ability for Gemini to make calls and send messages without unlocking your phone. Without this new feature enabled, when you attempt to make a call or send a message through Gemini, you’ll be greeted with the “Please unlock to continue” warning.

Hands-Free Calling and Messaging

With the feature enabled, Gemini can complete the request without you having to unlock your phone’s screen. This is a long-awaited feature because, as it stands, it’s impossible to enjoy true hands-free calling and messaging. Many times, when driving, I’ve needed to place a call but found myself having to pull over to unlock my phone before the call or message could be placed.

Rollout and Availability

The rollout is happening slowly, so you might not have the feature added yet. It’s hard to predict when your phone will receive the new feature because Google isn’t exactly forthcoming on these rollout dates. On my Pixel 9 Pro with Android 15 and the Google app version 15.47.29, the feature is still missing. Turns out the feature is available for version 15.48 of the Google app; until you have that version on the phone, you’ll still have to unlock your device before placing a call or sending a message using Gemini.

Checking for the Feature

You can check to see if you have the feature by opening Gemini, clicking your profile photo, tapping Settings > “Gemini on Lock Screen.” If you see an entry for “Make calls and send messages without unlocking,” congratulations! You have the feature. Otherwise, be patient: the latest version of the Google app eventually will arrive on your phone, and you’ll be able to place calls and send messages, via Gemini, without having to first unlock the device.

Future Features

Once Google sees how welcome this new feature is, perhaps the company will add other functions — like playing music from Gemini or starting Maps navigation — without the user having to first unlock the device. Until those features come to fruition, Android cannot claim true hands-free usage.

Conclusion

The addition of hands-free calling and messaging to Gemini is a significant step forward for the Google Assistant replacement. While the rollout is slow, it’s clear that Google is committed to making Gemini a more powerful and useful tool. With this feature, Android users can finally enjoy true hands-free usage, and we can only hope that Google will continue to add more features to make Gemini an even more compelling alternative to Google Assistant.

FAQs

Q: Will I receive the new feature automatically?
A: No, the rollout is happening slowly, and it’s hard to predict when your phone will receive the new feature.

Q: How do I check if I have the feature?
A: You can check by opening Gemini, clicking your profile photo, tapping Settings > “Gemini on Lock Screen.” If you see an entry for “Make calls and send messages without unlocking,” you have the feature.

Q: When will I get the latest version of the Google app?
A: The rollout of the latest version of the Google app is also slow, so it’s hard to predict when you’ll receive it. You can check for updates in the Google Play Store.

Q: What other features will be added to Gemini in the future?
A: Google has not announced any specific features, but it’s possible that the company will add other functions, such as playing music or starting Maps navigation, without the user having to first unlock the device.

How to Create a Web App in 2024 with React, tRPC, and Express

Building Fullstack Web Apps with TER: Prioritizing Developer-Friendly Solutions

The Vision Behind TER

TER was built with one main purpose: to make developers’ lives easier. It offers a straightforward configuration, cutting-edge tools, and an emphasis on speed, helping developers stay productive and focused on building rather than debugging or configuring.

What Sets TER Apart from Other Stacks?

Stacks like T3 often rely on frameworks like Next.js for handling frontend and backend needs. While effective, this approach isn’t ideal for every scenario. TER intentionally excludes Next.js, opting instead for a model where the frontend can be compiled into static files. These files can then be hosted on object storage solutions like AWS S3, making deployment simpler and more flexible.

Key Benefits of This Approach

  • Versatile Deployment Options: Frontend assets can be hosted on any static hosting platform, avoiding the need for specialized setups.
  • Designed for Web Applications: Unlike stacks optimized for SEO-heavy websites, TER is geared towards web apps like dashboards, internal tools, or applications where SEO isn’t a key concern. This allows developers to focus on functionality and user experience.

Why These Technologies?

The TER stack was curated to deliver a smooth, fullstack development experience while maintaining type safety and flexibility. Each component was chosen for its ability to simplify development without compromising power:

  • tRPC: Facilitates end-to-end type-safe communication between the frontend and backend, eliminating mismatched schemas and redundant validations.
  • Express: This lightweight framework for Node.js provides essential backend functionality with minimal overhead, giving developers complete control.
  • React: A powerful library for building interactive UIs, React complements TER’s focus on creating dynamic, user-friendly applications.
  • Drizzle ORM: Ensures robust, type-safe database interactions. Its integration with TypeScript aligns perfectly with TER’s philosophy of building reliable, maintainable applications.

Highlights of TER: A Comprehensive Toolkit for Fullstack Web Development

TER goes beyond being a simple collection of libraries. It includes essential features that address the common needs of modern web applications, from secure authentication to seamless data management.

  • Secure JWT Authentication with HttpOnly Cookies: Ensures stateless and secure authentication by using cookies inaccessible to JavaScript, significantly reducing the risk of XSS attacks.
  • External API Integration Examples: Demonstrates how to fetch data from APIs like Beers, Users, Movies, and Albums, making it easier to integrate third-party services into your application.
  • Health Monitoring Endpoint: A built-in /health endpoint allows for straightforward server status checks, critical for production environments.
  • Customizable Table Row Display: Lets users control how many rows are visible in data tables, enhancing usability for applications dealing with large datasets.

Why Drizzle Stands Out as an ORM

While Prisma has been a long-time favorite for its user-friendly design, scalability challenges in serverless environments pushed me to explore alternatives. The lack of optimized SQL queries and the overhead of the Prisma Rust query engine (which adds around 15MB to Lambda functions) became significant bottlenecks.

Drizzle emerged as a game-changer. Its TypeScript-first approach and SQL-level optimizations make it a better fit for TER, especially for serverless architectures. It’s leaner and more efficient, resolving connection pool issues without requiring paid add-ons like Prisma Accelerate.

Simplifying Development with npm Workspaces in a Monorepo

Past attempts to manage shared code between the frontend and backend using Git submodules or private npm packages proved cumbersome:

  • Code Sharing Issues: Submodules complicated workflows with multiple repos in the same IDE, while npm packages moved the source of truth away from the main repository.
  • Versioning Challenges: Aligning versions across repos was tedious with submodules, and tracking changes in npm packages was often confusing.
  • Slow Development Workflow: A single feature required multiple pull requests across repos, adding unnecessary delays.

Switching to npm Workspaces centralized all code in one monorepo, simplifying version control, reducing PR overhead, and enabling seamless code sharing—all while keeping GitHub as the source of truth.

Developer-Centric Features in TER

TER is engineered to maximize developer productivity through intelligent choices:

  • Type Safety Across the Stack: The combination of TypeScript and tRPC eliminates data type mismatches, ensuring smooth communication between frontend and backend.
  • Lightning-Fast Development with Vite: Vite’s hot module replacement significantly speeds up frontend iterations.
  • Extensibility at Its Core: The structure of TER allows developers to easily add routes, middleware, and components, offering a flexible starting point for any project.

Where TER Shines

TER’s tailored design makes it perfect for fullstack apps where SEO isn’t a primary concern:

  • Internal Dashboards: Build administrative tools, analytics panels, or reporting dashboards.
  • Data-Driven Applications: Handle large datasets efficiently with features like dynamic tables and robust API integrations.
  • Rapid Prototyping: Ideal for MVPs or proofs of concept that require quick iteration without compromising quality.

Installation Made Simple

Setting up TER is designed to be as hassle-free as possible:

  1. Create a Database: Set up a Postgres database named ter.
  2. Configure Environment Variables: Update env.ts with your database credentials.
  3. Install Dependencies: Run npm i from the root directory.
  4. Apply Migrations: Use npm run push to set up the database schema.
  5. Seed Data: Populate sample data with npm run seed.
  6. Launch the App: Start the development server with npm run dev to run both the backend and frontend.

Explore TER

TER isn’t just another stack—it’s a thoughtfully crafted toolkit for developers who value efficiency and flexibility. By taking a different path from conventional stacks like T3, TER excels in scenarios where SEO is not a priority, focusing on high-performance web applications.

Check out the project here: https://github.com/alan345/TER

If you find TER helpful, consider starring the repo to support its growth! 🚀

Evolv’s Security Scanners Falsely Hyped

Artificial Intelligence (AI) Companies and Public Safety: A Deadly Combination

Artificial intelligence (AI) companies are no strangers to embellishing their products’ capabilities — and when it comes to public safety, the consequences can be deadly.

Evolv Technologies’ False Claims

Evolv Technologies claimed its AI-powered security scanners would improve public safety by detecting weapons in a variety of places, including in schools. However, the Federal Trade Commission (FTC) found otherwise.

FTC Settlement Order

Last week, the FTC proposed a settlement order with Evolv Technologies over how the company oversold the ability of its AI-powered security system to detect weapons and "ignore harmless personal items" — especially in school settings. Evolv’s touch-free systems scan people as they enter a space without requiring a manual search of bags or pockets.

FTC Chair’s Statement

In an X post, FTC chair Lina Khan said Evolv has "falsely hyped" its AI weapons detection systems to school districts that paid the company "millions" to use the technology in schools. She noted that the scanners set off alarms for harmless items such as water bottles and binders while failing to identify weapons.

The Scanners’ Performance

Since going public in 2021, Evolv’s scanners have been popping up at sporting events, theme parks, schools, airports, subway stations, and even film festivals as an alternative to traditional metal detectors and other forms of security checkpoints. The Massachusetts-based company has advertised its "advanced AI-powered security scanning systems" as a viable solution to address public safety concerns, including school shootings.

Legal and Regulatory Scrutiny

In 2023, five law firms announced investigations into Evolv technologies for possible violations of securities law, claiming that "Evolv misled investors over the capabilities of its weapons detectors." Moreover, Evolv’s shareholders filed a class-action suit against the company, arguing that the company’s marketing claims overstated the effectiveness of the technology.

NYC Subway Pilot Program

Despite this legal, regulatory, and internal turmoil, New York City Mayor Eric Adams launched a three-month pilot program that would deploy the AI scanners in NYC subways. However, the scanners had been deployed previously in Jacobi Medical Center, where they reportedly triggered a huge number of false positives.

Conclusion

The FTC has been clear that claims about technology, including artificial intelligence, need to be backed up, and that is especially important when these claims involve the safety of children. The proposed settlement order would prevent Evolv from making unsupported claims about its products’ ability to detect weapons by using artificial intelligence. It would also require Evolv to notify certain K-12 school customers that they can opt to cancel contracts signed between April 1, 2022, to June 30, 2023.

Frequently Asked Questions

Q: What is the purpose of the FTC’s proposed settlement order with Evolv Technologies?
A: The proposed settlement order would prevent Evolv from making unsupported claims about its products’ ability to detect weapons by using artificial intelligence and would require Evolv to notify certain K-12 school customers that they can opt to cancel contracts signed between April 1, 2022, to June 30, 2023.

Q: What are the implications of Evolv’s false claims for public safety?
A: The implications are significant, as the company’s false claims may have led to the implementation of ineffective security measures, potentially putting the public at risk.

Q: What is the takeaway from this story?
A: The takeaway is that AI companies must be transparent about their products’ capabilities and limitations to ensure public safety.

Make a Movie with AI

0

Support and Community

Get Exclusive Content and Discounts

If you’re looking for a way to support me and get access to exclusive content, I invite you to join my Patreon community. By becoming a patron, you’ll gain access to my workflows, which can help you improve your own creative projects. Use the link below to sign up and start supporting me today!

Patreon: https://www.patreon.com/oliviotutorials

Get a 1€ Discount on Qwertee

In addition to supporting my work, you can also get a 1€ discount on your next purchase at Qwertee.com. Just use the code "Olivio" at checkout to receive your discount. This is a great way to score some amazing deals on unique and creative designs.

Join My Communities

Stay connected with me and other creatives by joining my Discord group and Facebook group. These communities are a great place to share your work, get feedback, and learn from others.

Discord: https://discord.gg/XKAk7GUzAW
Facebook: https://www.facebook.com/groups/theairevolution

Source

Frequently Asked Questions

Q: What is Patreon and how does it work?
A: Patreon is a platform that allows creators to earn money from their fans and supporters. By becoming a patron, you’ll receive exclusive content and rewards in exchange for your support.

Q: What kind of content can I expect to get on Patreon?
A: As a patron, you’ll gain access to my workflows, which can help you improve your own creative projects. This may include tutorials, tips, and behind-the-scenes content.

Q: How much does it cost to join Patreon?
A: The cost to join Patreon varies depending on the tier you choose. You can select from a variety of options, from $1 to $50 per month.

Q: Can I still access the content if I’m not a patron?
A: Yes, you can still access my public content, but becoming a patron will give you exclusive access to premium content and rewards.

AI Agents to Dominate Enterprise Deployment by 2025

Enterprise Use of AI Agents on the Rise

According to Deloitte’s Global 2025 Predictions Report, enterprise use of AI agents is on the rise, with 25% of enterprises using generative AI forecast to deploy AI agents in 2025, growing to 50% by 2027.

Autonomous Generative AI Agents

Deloitte defines autonomous generative AI agents, also known as agentic AI, as software solutions that can complete complex tasks and meet objectives with little or no human supervision.

Interesting Predictions

  • Women’s Adoption Gap in Gen AI Usage is Closing Quickly: By 2025, women’s experimentation and usage of GenAI are projected to meet or exceed that of men, but tech companies still should improve trust, representation in training models, and diversity in the AI workforce.
  • Gen AI is Driving Data Center Energy Consumption Surge: Electricity consumption by global data centers is forecasted to double to 4% (1,065 terawatt-hours) by 2030 as power-intensive Gen AI consumption grows faster than other uses and applications.
  • Gen AI is Set to Make Devices Smarter: In 2025, the share of shipped Gen AI-enabled smartphones could exceed 30%, in addition to about 50% of laptops with local Gen AI processing capabilities.

Characteristics and Capabilities of Agentic AI

  • Built on Foundation Models: Foundation models like LLMs enable agentic AI to reason, analyze, and adapt to complex and unpredictable workflows, making them more flexible than RPA and expert systems.
  • Acts Autonomously: While the degree of autonomy varies, agentic AI can be trained to plan and execute complex tasks largely on its own.
  • Senses the Environment: Agentic AI can perceive the environment, process information, and understand the context of the tasks it is given.
  • Uses Tools: Agentic AI interacts with tools and systems to complete tasks, such as software, enterprise applications, and the internet.
  • Orchestrates: Agentic AI can direct the participation of other systems and bots to complete a task.
  • Accesses Memory: Agentic AI can access short-term memory to maintain context while performing a specific task, and long-term memory to learn and improve from experience.

Conclusion

As agentic AI continues to evolve, it is essential for businesses to understand its capabilities and limitations. With the ability to complete complex tasks and meet objectives with little or no human supervision, agentic AI has the potential to revolutionize the way we work. However, it is crucial to address the challenges and concerns surrounding its adoption, such as ensuring trust, representation, and diversity in the AI workforce.

FAQs

Q: What is agentic AI?
A: Agentic AI is software that can complete complex tasks and meet objectives with little or no human supervision.

Q: What are the characteristics of agentic AI?
A: Agentic AI is built on foundation models, acts autonomously, senses the environment, uses tools, orchestrates, and accesses memory.

Q: What are the benefits of agentic AI?
A: Agentic AI has the potential to revolutionize the way we work, completing complex tasks and meeting objectives with little or no human supervision.

Q: What are the challenges of agentic AI?
A: Ensuring trust, representation, and diversity in the AI workforce are crucial challenges to address when adopting agentic AI.

Document Hubs for Unified Data Access and Governance

The Rise of Document Hubs: Simplifying Data Discovery and Collaboration

The way organizations manage and share data can make or break their ability to collaborate and make informed decisions. Inconsistent data definitions and fragmented documentation often create barriers between teams, limiting efficiency and trust in decision-making.

Alation, a data intelligence company, is aiming to address this challenge with the release of Document Hubs – a centralized platform designed to simplify data discovery and foster collaboration within organizations. The platform is designed to align teams on a shared data language, providing consistent access to essential documentation like glossaries, data products, and AI insights.

A Fragmented Landscape

In a 2023 McKinsey study, 80% of businesses struggle with isolated data systems, each using unique standards and practices. This disconnect often hampers data accessibility and consistency, stifling innovation and collaboration. As organizations grow, these challenges worsen. Data becomes scattered, definitions become disconnected, and unclear documentation makes it increasingly difficult to align teams and extract meaningful insights.

Introducing Document Hubs

According to Alation, Document Hubs offers a unified solution for businesses to organize and access knowledge and data semantics specific to their organization’s needs. It simplifies access to role-specific information, eliminating the need to search across multiple systems.

Key Features

  • Create custom document types, such as business processes, project documentation, policies, metric definitions, or anything else that suits your needs
  • Create multiple document hubs for different purposes, allowing you to have documents tailored to specific object types
  • Organize documentation in a structured hierarchy, enabling users to easily find specific information
  • Nested folders and sub-documents improve content management
  • Connect data assets, making it easier for users to find relevant information
  • Custom templates, metadata, and tags improve searchability, supporting faster, data-driven decisions

Real-World Implementation

For Burns & McDonnell, Alation’s Document Hubs serves as a knowledge semantic layer for data. The Kansas City-based engineering consultancy has used the platform to centralize critical documentation, streamline data management, and enhance accessibility for teams across the organization.

Conclusion

The release of Document Hubs builds on Alation’s momentum, expanding its governance capabilities to include comprehensive knowledge and semantic management across the enterprise. By providing a centralized platform for data discovery and collaboration, Alation is transforming data from a fragmented resource into a cohesive, strategic asset.

FAQs

Q: What is Document Hubs?
A: Document Hubs is a centralized platform designed to simplify data discovery and foster collaboration within organizations.

Q: What are the key features of Document Hubs?
A: Key features include custom document types, multiple document hubs, structured hierarchy, nested folders, and custom templates.

Q: How can Document Hubs help businesses?
A: Document Hubs simplifies access to role-specific information, eliminates the need to search across multiple systems, and supports faster, data-driven decisions.

Q: Who can benefit from Document Hubs?
A: Any organization seeking to simplify data discovery and foster collaboration can benefit from Document Hubs.

NVIDIA Contributes NVIDIA GB200 NVL72 Designs to Open Compute Project

0

Here is the organized content:

NVIDIA Open-Source Initiatives

NVIDIA has a rich history of open-source initiatives. NVIDIA engineers have released over 900 software projects on GitHub and have open-sourced essential components of the AI software stack. The NVIDIA Triton Inference Server, for example, is now integrated into all major cloud service providers to serve AI models in production. Additionally, NVIDIA engineers are actively involved in numerous open-source foundations and standards bodies, including the Linux Foundation, the Python Software Foundation, and the PyTorch Foundation.

Meeting Data Center Compute Demands

The compute power required to train autoregressive transformer models has exploded, growing by a staggering 20,000x over the last 5 years. Meta’s Llama 3.1 405B model, launched earlier this year, required 38 billion petaflops of accelerated compute to train, 50x more than the Llama 2 70B model launched only a year earlier. Training and serving these large models cannot be managed on a single GPU; rather, they must be parallelized across massive GPU clusters.

Importance of Multi-GPU Interconnect

One common challenge that arises with model parallelism is the high volume of GPU-to-GPU communication. Tensor parallel GPU communication patterns highlight just how interconnected these GPUs are. For example, with AllReduce, every GPU has to send the results of its calculation to every other GPU at every layer of the neural network before the final model output is determined. Any latency during these communications can lead to significant inefficiencies, with GPUs left idle, waiting for the communication protocols to complete. This reduces overall system efficiency and increases the total cost of ownership (TCO).

NVLink Cartridges

To enable high-speed communication between all 72 NVIDIA Blackwell GPUs in the NVLink domain, we implemented a novel design featuring four NVLink cartridges mounted vertically at the rear of the rack. These cartridges accommodate over 5,000 active copper cables, delivering an impressive aggregate All-to-All bandwidth of 130 TB/s and 260 TB/s AllReduce bandwidth.

Liquid Cooling Manifolds and Floating Blind Mates

To efficiently manage the 120 KW cooling capacity required for the rack, we’ve implemented direct liquid cooling techniques. Building upon existing designs, we’ve introduced two key innovations. First, we developed an enhanced Blind Mate Liquid Cooling Manifold design, capable of delivering efficient cooling. Second, we created a novel Floating Blind Mate Tray connection, which effectively distributes coolant to both compute and switch trays, significantly improving the ability of the liquid quick disconnects to align and reliably mate in the rack.

Compute and Switch Tray Mechanical Form Factors

To accommodate the high compute density of the rack, we introduced 1RU liquid-cooled compute and switch tray form factors. We also developed a new, denser DC-SCM (Data Center Secure Control Module) design that’s 10% smaller than the current standard. In addition, we implemented a narrower bus bar connector to maximize available rear panel space. These modifications optimize space utilization while maintaining performance.

New Joint NVIDIA GB200 NVL72 Reference Architecture

At OCP, NVIDIA also announced a new joint GB200 NVL72 reference architecture with Vertiv, a leader in power and cooling technologies and expert in designing, building, and servicing high compute density data centers. This new reference architecture will significantly reduce implementation time for CSPs and data centers deploying the NVIDIA Blackwell platform.

Conclusion

The NVIDIA GB200 NVL72 design represents a significant milestone in the evolution of modern high compute density data centers. By addressing the pressing challenges of training and serving growing AI models and high GPU-to-GPU communication, this contribution accelerates the adoption of energy-efficient high compute density platforms in the data center while reinforcing the importance of collaboration within the open ecosystem. We’re excited to see how the OCP community will leverage and build on top of the GB200 NVL72 design contributions.

FAQs

Q: What is the NVLink domain size and speed in the GB200 NVL72 design?
A: The NVLink domain can now support up to 72 NVIDIA Blackwell GPUs, with a communication speed of 1.8 TB/s per GPU, 36x faster than state-of-the-art 400 Gbps Ethernet standards.

Q: What is the aggregate All-to-All bandwidth of the NVLink domain in the GB200 NVL72 design?
A: The aggregate All-to-All bandwidth is 260 TB/s.

Q: What is the thermal management solution used in the GB200 NVL72 design?
A: The GB200 NVL72 design uses direct liquid cooling techniques, including an enhanced Blind Mate Liquid Cooling Manifold design and a novel Floating Blind Mate Tray connection.

Q: What is the compute and switch tray mechanical form factor in the GB200 NVL72 design?
A: The GB200 NVL72 design features 1RU liquid-cooled compute and switch tray form factors.

Q: Who is the partner that NVIDIA collaborated with to create the new joint GB200 NVL72 reference architecture?
A: NVIDIA collaborated with Vertiv, a leader in power and cooling technologies and expert in designing, building, and servicing high compute density data centers.

5 Business Tactics for a Successful AI Transformation

Aiming for a Successful AI Transformation: 5 Ways to Ensure Your Organization Reaps the Rewards

Aiming for a successful AI transformation is great, but if you can’t lead the initiative effectively, you won’t deliver the results the business demands.

1. Partner with your peers

Gabriela Vogel, senior director analyst in the Executive Leadership of Digital Business practice at research firm Gartner, said digital leaders must focus urgently on exploiting emerging technology effectively. "CIOs who don’t understand the focus on value — and make promises about AI without thinking about what they are getting involved with — might not stay at the top. They will lose power, they will lose inference, and potentially even lose their jobs."

2. Create a working group

James Fleming, CIO at Francis Crick Institute, said the digital leader’s role in AI is to provide oversight within the IT department and across the business. "There’s got to be a degree of leading from the front and making it OK for your team to think about these questions and become experts in them, and then you’ve got to guide the rest of the organization along that journey."

3. Optimise your resources

Bruno Marie-Rose, chief information and technology officer of the Paris 2024 Organising Committee for the Olympic and Paralympic Games, said the key to leading AI transformations is turning new personal habits into business benefits. "Maturity is crucial for the International Organizing Committee, having an approach where you can say, ‘I’m four years ahead of the Opening Ceremony and need to progress, what do you advise?’"

4. Reduce people’s fears

Ollie Wildeman, vice president of customer services at travel specialist Big Bus Tours, said executives who explore AI will discover three types of people are worried. "It will be the people on the front line who think their jobs will be replaced, it will be the people who manage those guys, and stakeholders will worry about the money you’re putting in and what will happen to the customer satisfaction scores in the long run."

5. Don’t poison the well

Jon Grainger, CTO at legal firm DWF, said successful AI transformations ensure the data that feeds IT systems is well-managed and trustworthy. "You can’t do stuff without your data being right. You can go much deeper — and be much more solid on your outcomes and what you’re trying to achieve — as your data gets better."

Conclusion

In conclusion, to ensure a successful AI transformation, organizations must partner with peers, create a working group, optimize resources, reduce people’s fears, and don’t poison the well. By following these five ways, organizations can reap the rewards of AI and stay ahead of the competition.

FAQs

Q: What is the role of the digital leader in AI?
A: The digital leader’s role in AI is to provide oversight within the IT department and across the business.

Q: How can organizations optimize resources with AI?
A: Organizations can optimize resources with AI by using data to help optimize the use of resources across the Olympics.

Q: How can executives reduce people’s fears about AI?
A: Executives can reduce people’s fears about AI by communicating the benefits of AI and providing training and support for those who will be impacted by the technology.

Q: What is the importance of data quality in AI?
A: Data quality is crucial in AI, as poor data can lead to incorrect results and undermine the effectiveness of the technology.