Home Blog Page 324

DeepSeek iOS App Sends Data Unencrypted

Security Experts Weigh In on DeepSeek’s Data Security Concerns

Experts Slam DeepSeek’s Use of Unencrypted HTTP Endpoints

According to a recent report, popular social media app DeepSeek has been found to be using unencrypted HTTP endpoints, allowing sensitive data to be transmitted without encryption. This has raised concerns among security experts, who are warning that this could pose a significant risk to national security.

"Bad Idea" Says Thomas Reed

Thomas Reed, staff product manager for Mac endpoint detection and response at security firm Huntress, and an expert in iOS security, has expressed his concerns about DeepSeek’s use of unencrypted HTTP endpoints. "ATS being disabled is generally a bad idea," he wrote in an online interview. "That essentially allows the app to communicate via insecure protocols, like HTTP. Apple does allow it, and I’m sure other apps probably do it, but they shouldn’t. There’s no good reason for this in this day and age."

Even with Encryption, Security Experts Unwilling to Trust DeepSeek

Reed also emphasized that even if the app were to secure its communications, he would still be unwilling to send sensitive data to a server that the government of China could access. "Even if they were to secure the communications, I’d still be extremely unwilling to send any remotely sensitive data that will end up on a server that the government of China could get access to," he said.

Others Less Concerned about Chinese Companies’ Access to Data

HD Moore, founder and CEO of runZero, took a different view, expressing less concern about ByteDance or other Chinese companies having access to data. "The unencrypted HTTP endpoints are inexcusable," he wrote. "You would expect the mobile app and their framework partners (ByteDance, Volcengine, etc) to hoover device data, just like anything else—but the HTTP endpoints expose data to anyone in the network path, not just the vendor and their partners."

Government Reactions

On Thursday, US lawmakers began pushing to immediately ban DeepSeek from all government devices, citing national security concerns that the Chinese Communist Party may have built a backdoor into the service to access Americans’ sensitive private data. If passed, DeepSeek could be banned within 60 days.

Conclusion

The use of unencrypted HTTP endpoints by DeepSeek has sparked a heated debate among security experts, with some expressing concern about the potential risks to national security. As the debate continues, it is clear that the security of user data is a top priority.

FAQs

Q: What are unencrypted HTTP endpoints?
A: Unencrypted HTTP endpoints refer to the use of unsecured communication protocols, such as HTTP, to transmit sensitive data.

Q: Why is this a security risk?
A: Unencrypted data can be intercepted and accessed by anyone in the network path, including malicious actors and governments.

Q: Is this a common practice among apps?
A: No, using unencrypted HTTP endpoints is not a common practice among reputable apps. Many apps use encryption to protect user data.

Q: What is being done to address these concerns?
A: US lawmakers are pushing to ban DeepSeek from all government devices, citing national security concerns. The app’s developers have not commented on the issue.

The VSCode One Looks Awesome

10 Creative & Open Source Portfolio Templates

Introduction

A portfolio is a crucial tool for any professional, allowing them to showcase their skills and experience to potential employers. While there are many portfolio templates available, not all of them are free or open source. In this article, we will explore 10 creative and open source portfolio templates that you can use to build your own portfolio.

1. Gridster

Gridster is a popular open source portfolio template that is designed to be highly customizable. It features a grid-based layout that can be easily modified to fit your needs. Gridster is built using HTML, CSS, and JavaScript, making it a great option for web developers.

2. Simple Portfolio

Simple Portfolio is a clean and minimalist template that is perfect for those who want to showcase their work in a straightforward way. The template is built using HTML and CSS, and is highly customizable.

3. Parallax Portfolio

Parallax Portfolio is a unique template that features a parallax scrolling effect. The template is built using HTML, CSS, and JavaScript, and is a great option for those who want to add some visual interest to their portfolio.

4. Masonry

Masonry is a popular open source portfolio template that is designed to be highly customizable. The template is built using HTML, CSS, and JavaScript, and features a masonry layout that is perfect for showcasing a variety of projects.

5. Flexbox Portfolio

Flexbox Portfolio is a template that is built using HTML, CSS, and JavaScript. The template features a flexible layout that can be easily modified to fit your needs.

6. Foundation

Foundation is a popular open source portfolio template that is designed to be highly customizable. The template is built using HTML, CSS, and JavaScript, and features a responsive design that looks great on any device.

7. Bootstrap

Bootstrap is a popular open source portfolio template that is designed to be highly customizable. The template is built using HTML, CSS, and JavaScript, and features a responsive design that looks great on any device.

8. Jekyll

Jekyll is a popular open source portfolio template that is designed to be highly customizable. The template is built using HTML, CSS, and JavaScript, and features a responsive design that looks great on any device.

9. Hugo

Hugo is a popular open source portfolio template that is designed to be highly customizable. The template is built using HTML, CSS, and JavaScript, and features a responsive design that looks great on any device.

10. Gatsby

Gatsby is a popular open source portfolio template that is designed to be highly customizable. The template is built using HTML, CSS, and JavaScript, and features a responsive design that looks great on any device.

Conclusion

These 10 creative and open source portfolio templates are a great way to showcase your skills and experience. Whether you’re a web developer, designer, or writer, these templates can help you build a professional-looking portfolio that will help you stand out in the job market.

FAQs

Q: Are these templates free?

A: Yes, all of the templates listed above are open source and free to use.

Q: Can I customize these templates?

A: Yes, all of the templates listed above are highly customizable, allowing you to modify the design and layout to fit your needs.

Q: Do these templates require coding knowledge?

A: Some of the templates listed above require coding knowledge, while others are more user-friendly. If you’re not a coder, you may want to start with one of the more user-friendly templates.

Q: Can I use these templates on my own website?

A: Yes, all of the templates listed above are designed to be used on your own website, allowing you to showcase your work and experience to potential employers.

Amazon to Spend $100bn in AI Drive

Amazon to Pump $100bn into AI Initiatives This Year

Amazon has announced that it will invest around $100bn in its artificial intelligence (AI) initiatives this year, as the e-commerce group shrugs off concerns about China’s DeepSeek and invests heavily in data infrastructure.

AI Investment to Focus on Amazon Web Services

The bulk of the investment will be in Amazon Web Services (AWS), which operates data centers and offers customers software tools. This move is seen as a strategic one, given the rapid growth of the AI industry and the increasing demand for cost-effective and efficient tools.

Jassy’s Vision for AI

Amazon’s chief executive, Andy Jassy, has stated that he sees "significant signals of demand" for AI services and products, and that the prospect of cheaper and more efficient tools will lead to more customer spending. He believes that companies will spend less per unit of infrastructure, but ultimately end up spending more in total.

Sales Growth and Revenue

Amazon’s sales growth in the fourth quarter was modest, with revenues rising 10% year on year to $187.8bn. However, the company expects net sales in the current quarter to come in between $151bn and $155.5bn, which is below forecasts.

Cost-Cutting Efforts

Jassy has overseen a cost-cutting effort in recent years, including streamlining logistics operations and reducing middle management. This has allowed the company to invest in data center capacity and expand its AI initiatives.

Conclusion

Amazon’s massive AI investment is a significant step towards dominating the rapidly growing industry, and its ability to attract top talent and secure large-scale deals will be crucial to its success. While the company’s sales growth may be modest, its investments in AI and data infrastructure will likely pay off in the long run.

FAQs

  • Q: What is Amazon’s AI investment strategy?
    A: Amazon will invest around $100bn in its AI initiatives this year, with a focus on Amazon Web Services (AWS).
  • Q: What is the purpose of Amazon’s AI investment?
    A: The purpose is to develop cost-effective and efficient AI tools, which will lead to increased customer spending and drive growth.
  • Q: How does Amazon’s cost-cutting effort affect its investment in AI?
    A: Amazon’s cost-cutting efforts have allowed it to invest in data center capacity and expand its AI initiatives.
  • Q: What is the outlook for Amazon’s sales growth?
    A: Amazon’s sales growth is expected to be modest, but its investments in AI and data infrastructure will likely pay off in the long run.

Amazon Reports Rise in Profits

0

Amazon Reports Strong Q4 Results, But Warns of Slowing Growth Ahead

Revenue and Profit Growth

Amazon reported strong revenue and profit growth in its latest quarter, with sales reaching $187.8 billion, a 10% increase from the same period last year. Profit rose 88% to $20 billion, roughly in line with Wall Street expectations.

Cloud Computing Growth

The company’s cloud computing business, which is a significant contributor to its revenue and profit, grew 19% to $28.8 billion, with operating profit of $10.6 billion, accounting for half of Amazon’s overall operating profit. This growth is particularly strong, given that its top cloud competitors, Microsoft and Alphabet, recently reported underwhelming results.

Artificial Intelligence (AI) and Capital Expenditures

Amazon’s investments in AI are paying off, with the company’s cloud business becoming a profit engine. The company told investors that it is making its AI system available to customers, allowing them to easily mix and match different AI tools. Amazon has also spent $26 billion building out data centers, warehouses, and other capital expenses in the quarter, bringing its annual total to over $77 billion. The company expects to continue this level of capital investment throughout 2025, potentially reaching over $100 billion this year.

North American Retail Business

Amazon’s North American retail business, which includes product sales and services like advertising and Prime memberships, grew 10% in the holiday shopping period, outpacing the retail industry’s overall growth. The company’s efforts to offer faster shipping and rework its logistics have helped improve its operating margin, which grew to over 8%, up from 2% two years ago.

Employee Count and Share Price

Amazon ended the year with 1,556,000 employees, a 2% increase. The company’s share price was down more than 4% in after-hours trading.

Conclusion

Amazon’s strong Q4 results are a testament to its profitable cloud business and efforts to improve its retail business. However, the company has warned investors to expect slower growth ahead, as its capital expenditures continue to increase. Despite this, Amazon’s strong financial position and growth prospects make it an attractive investment opportunity.

FAQs

Q: What are Amazon’s expectations for the current quarter?
A: Amazon expects sales to grow between 5% and 9% in the current quarter, with operating profit potentially lower than last year.

Q: How is Amazon’s cloud business performing?
A: Amazon’s cloud business grew 19% to $28.8 billion, with operating profit of $10.6 billion, accounting for half of Amazon’s overall operating profit.

Q: What is Amazon’s strategy for artificial intelligence (AI)?
A: Amazon is making its AI system available to customers, allowing them to easily mix and match different AI tools.

Q: How is Amazon’s retail business performing?
A: Amazon’s North American retail business grew 10% in the holiday shopping period, with operating margin growing to over 8%.

When the Earth Talks, AI Listens

AI Built for Speech Is Now Decoding the Language of Earthquakes

AI, originally designed for speech recognition, has been repurposed to analyze seismic signals from Hawaii’s 2018 Kīlauea volcano collapse. The findings, published in Nature Communications, suggest that faults emit distinct signals as they shift — patterns that AI can now track in real-time.

Seismic records are acoustic measurements of waves passing through the solid Earth, said Christopher Johnson, one of the study’s lead researchers. “From a signal processing perspective, many similar techniques are applied for both audio and seismic waveform analysis.”

The AI model was tested using data from the 2018 collapse of Hawaii’s Kīlauea caldera, which triggered months of earthquakes and reshaped the volcanic landscape. Big earthquakes don’t just shake the ground — they upend economies. In the past five years, quakes in Japan, Turkey, and California have caused tens of billions of dollars in damage and displaced millions of people.

How AI Was Trained to Listen to the Earth

Unlike previous machine learning models that required manually labeled training data, the researchers used a self-supervised learning approach to train Wav2Vec-2.0. The model was pre-trained on continuous seismic waveforms and then fine-tuned using real-world data from Kīlauea’s collapse sequence. NVIDIA accelerated computing played a crucial role in processing vast amounts of seismic waveform data in parallel.

What’s Still Missing: Can AI Predict Earthquakes?

While the AI showed promise in tracking real-time fault shifts, it was less effective at forecasting future displacement. Attempts to train the model for near-future predictions — essentially, asking it to anticipate a slip event before it happens — yielded inconclusive results. “We need to expand the training data to include continuous data from other seismic networks that contain more variations in naturally occurring and anthropogenic signals,” he explained.

A Step Toward Smarter Seismic Monitoring

Despite the challenges in forecasting, the results mark an intriguing advancement in earthquake research. This study suggests that AI models designed for speech recognition may be uniquely suited to interpreting the intricate, shifting signals faults generate over time. “This research, as applied to tectonic fault systems, is still in its infancy,” Johnson said. “The study is more analogous to data from laboratory experiments than large earthquake fault zones, which have much longer recurrence intervals. Extending these efforts to real-world forecasting will require further model development with physics-based constraints.”

Conclusion

The study marks an important step toward understanding how faults behave before a slip event. While the AI model is not yet capable of predicting earthquakes, it has the potential to greatly improve our ability to track fault movements in real-time. By expanding the training data and refining the model, scientists may be able to create a more accurate and reliable tool for predicting earthquake activity.

Frequently Asked Questions

Q: Can AI really predict earthquakes?
A: While the AI model showed promise in tracking real-time fault shifts, it was less effective at forecasting future displacement. Attempts to train the model for near-future predictions yielded inconclusive results.

Q: What is the current state of AI in earthquake research?
A: The study is still in its infancy, but it marks an important advancement in earthquake research. The AI model has the potential to greatly improve our ability to track fault movements in real-time.

Q: How was the AI model trained?
A: The researchers used a self-supervised learning approach to train Wav2Vec-2.0. The model was pre-trained on continuous seismic waveforms and then fine-tuned using real-world data from Kīlauea’s collapse sequence.

Q: What role did NVIDIA play in the study?
A: NVIDIA accelerated computing played a crucial role in processing vast amounts of seismic waveform data in parallel. High-performance NVIDIA GPUs accelerated training, enabling the AI to efficiently extract meaningful patterns from continuous seismic signals.

AI Reveals Thought Process

0

OpenAI Updates Chain of Thought for o3-mini AI Model

In Response to Rival’s Pressure, OpenAI Changes the Way o3-mini Communicates its Step-by-Step "Thought" Process

OpenAI, a leading AI research organization, is modifying the way its newest AI model, o3-mini, communicates its step-by-step "thought" process in response to pressure from rivals, including Chinese AI company DeepSeek. The company is introducing an updated "chain of thought" that shows more of the model’s "reasoning" steps and how it arrives at answers to questions.

Updated Chain of Thought for ChatGPT Users

The updated chain of thought will be available to free and paid users of ChatGPT, OpenAI’s AI-powered chatbot platform. Subscribers to premium ChatGPT plans who use o3-mini in the "high reasoning" configuration will also see this updated readout. According to OpenAI, the updated chain of thought will make it easier for users to understand how the model thinks, providing more clarity and confidence in its responses.

How the Update Works

The update introduces an additional post-processing step where the model reviews its raw chain of thought, removing any unsafe content and simplifying complex ideas. This step also enables non-English users to receive the chain of thought in their native language, making it more accessible and friendly.

Rationale Behind the Update

OpenAI’s decision to update the chain of thought is partly due to competitive reasons. DeepSeek’s R1 model, a "reasoning" model similar to o3-mini, reveals its full thought process, which many AI researchers argue is the preferred approach. The reasoning steps deliver a better user experience in certain situations, helping to indicate when the model might be on the right or wrong track.

Background on o3-mini

o3-mini is a "reasoning" model that thoroughly fact-checks itself before giving out results. This approach helps the model avoid some of the pitfalls that normally trip up models. The trade-off is that reasoning models take a little longer to arrive at solutions – typically seconds to minutes longer.

Reaction to the Update

Noam Brown, a researcher, tweeted about the update, saying, "When we briefed people on 🍓 before o1-preview’s release, seeing the CoT live was usually the ‘aha’ moment for them that made it clear this was going to be a big deal. These aren’t the raw CoTs but it’s a big step closer and I’m glad we can share that experience with the world."

Frequently Asked Questions

Q: What is the purpose of the updated chain of thought?
A: The updated chain of thought is designed to make it easier for users to understand how the model thinks, providing more clarity and confidence in its responses.

Q: Who will benefit from the update?
A: Free and paid users of ChatGPT, as well as subscribers to premium ChatGPT plans who use o3-mini in the "high reasoning" configuration.

Q: How does the update address competitive concerns?
A: The update addresses competitive concerns by providing more transparency into the model’s thought process, similar to DeepSeek’s R1 model.

MIND BLOWING!

0

Revolutionizing Video Creation: ByteDance’s OmniHuman AI

Introduction

ByteDance, the parent company of popular social media platforms TikTok and Douyin, has recently released a groundbreaking AI technology that is poised to change the game for video creation. The OmniHuman AI is a revolutionary tool that uses artificial intelligence to generate human-like videos, opening up new possibilities for content creators and marketers.

What is OmniHuman AI?

The OmniHuman AI is a deep learning-based technology that uses advanced algorithms to create realistic human-like videos. The AI is trained on a vast dataset of videos and can generate content that is indistinguishable from real human-made videos. The technology uses a combination of computer vision, natural language processing, and machine learning to create videos that are both realistic and engaging.

How Does it Work?

The OmniHuman AI works by analyzing a dataset of videos and identifying patterns and trends. It then uses this information to generate new videos that are similar in style and content to the original videos. The AI can generate videos in various formats, including 2D and 3D, and can even add special effects and music to make the videos more engaging.

Applications of OmniHuman AI

The OmniHuman AI has a wide range of applications in various industries, including:

  • Content Creation: The AI can be used to generate high-quality content for social media platforms, YouTube, and other online channels.
  • Marketing: The AI can be used to create engaging marketing videos that can be used to promote products and services.
  • Education: The AI can be used to create interactive educational videos that can be used to teach students new skills and concepts.
  • Healthcare: The AI can be used to create videos that can be used to educate patients about medical conditions and treatments.

Benefits of OmniHuman AI

The OmniHuman AI offers several benefits, including:

  • Increased Efficiency: The AI can generate videos much faster than human creators, making it an ideal solution for businesses that need to produce high-quality content quickly.
  • Cost Savings: The AI can reduce the cost of video production by eliminating the need for human creators and reducing the need for expensive equipment and software.
  • Improved Quality: The AI can generate high-quality videos that are indistinguishable from real human-made videos.

Conclusion

The OmniHuman AI is a groundbreaking technology that is poised to revolutionize the video creation industry. With its ability to generate human-like videos, the AI has the potential to change the way we create and consume content. Whether you’re a content creator, marketer, or educator, the OmniHuman AI is an exciting technology that is definitely worth exploring.

FAQs

Q: What is the OmniHuman AI?
A: The OmniHuman AI is a deep learning-based technology that uses advanced algorithms to generate human-like videos.

Q: How does the OmniHuman AI work?
A: The OmniHuman AI works by analyzing a dataset of videos and identifying patterns and trends. It then uses this information to generate new videos that are similar in style and content to the original videos.

Q: What are the applications of the OmniHuman AI?
A: The OmniHuman AI has a wide range of applications in various industries, including content creation, marketing, education, and healthcare.

Q: What are the benefits of the OmniHuman AI?
A: The OmniHuman AI offers several benefits, including increased efficiency, cost savings, and improved quality.

AI Powers the Perfect Super Bowl Experience

0

Preparing for the Future of Work: An Expert Insights

The Need for Adaptation

In an increasingly fast-paced and ever-changing job market, it’s crucial for professionals to be prepared for the future of work. In an exclusive interview, our expert shares valuable insights on what it takes to thrive in this new landscape.

Embracing the Gig Economy

The gig economy has become a reality, with more and more professionals opting for flexible work arrangements. According to our expert, "The gig economy is here to stay, and it’s essential to be adaptable and open to new opportunities."

Building a Strong Professional Network

In today’s digital age, networking has become more important than ever. Our expert emphasizes the importance of building a strong professional network, stating, "Networking is key to staying ahead in the game. It’s not just about who you know, but also about who knows you."

Staying Relevant in the Job Market

With the rise of automation and AI, it’s more important than ever to stay relevant in the job market. Our expert advises, "Stay curious, stay learning, and stay open to new experiences. The future of work is about being agile and adaptable."

The Role of Soft Skills in the Future of Work

Soft skills such as communication, empathy, and problem-solving are becoming increasingly valuable in the job market. Our expert notes, "Soft skills are essential for success in the future of work. They help you navigate complex situations and build strong relationships."

Conclusion

In conclusion, preparing for the future of work requires adaptability, a strong professional network, and a focus on soft skills. By embracing the gig economy, staying relevant in the job market, and building strong relationships, professionals can thrive in this new landscape.

FAQs

Q: What is the gig economy, and how does it impact the job market?

A: The gig economy refers to a labor market characterized by short-term, flexible, and often freelance work arrangements. It has significantly impacted the traditional 9-to-5 job structure, with more professionals opting for flexible work arrangements.

Q: How can I build a strong professional network?

A: Building a strong professional network requires consistent effort and engagement. Attend industry events, join online communities, and connect with people on LinkedIn. Most importantly, be authentic and genuine in your interactions.

Q: What are the most important soft skills for the future of work?

A: According to our expert, the most important soft skills for the future of work are communication, empathy, and problem-solving. These skills help you navigate complex situations and build strong relationships.

AMC Upping Price for A-List Stubs Subscription

0

AMC Raises Prices for Stubs A-List Subscription

Following AMC’s recent changes to its Stubs program, the theater chain is now planning to raise the subscription service’s prices. Beginning May 7th, a membership to Stub’s A-List tier will increase from $24.95 / month to $27.99 / month.

New Benefits to Sweeten the Deal

Despite the price hike, AMC is trying to make the subscription more attractive to customers. As part of the changes, Stubs A-List members will be able to watch four movies a week at any AMC theater nationwide, up from the current three movies a week. Additionally, AMC is lowering the age requirement for a Stubs A-List membership to 13, aiming to court more young film buffs.

CEO’s Defense of the Price Increase

AMC CEO Adam Aron defended the price hike, stating that “A-List is still an incredible bargain.” He noted that, even with the increase, the cost of an A-List membership will often be less than seeing two movies per month as a non-member, especially for those who see movies in premium formats and/or buy their tickets online.

Conclusion

The changes to AMC’s Stubs program aim to strike a balance between providing value to customers and generating revenue. While the price hike may not be welcome news for some subscribers, the addition of new benefits and a lower age requirement may entice others to sign up. As the theater chain continues to evolve, it will be interesting to see how these changes affect its loyal customer base.

FAQs

Q: When will the price increase take effect?

A: The price increase will take effect on May 7th.

Q: What is changing about the Stubs A-List program?

A: The number of movies subscribers can watch weekly is increasing from three to four, and the age requirement for a Stubs A-List membership is being lowered to 13.

Q: Is this the first price increase for AMC’s Stubs program?

A: According to AMC CEO Adam Aron, this is the service’s first price hike in years.

Q: Is the price increase still a good deal?

A: According to Aron, the cost of an A-List membership will often be less than seeing two movies per month as a non-member, especially for those who see movies in premium formats and/or buy their tickets online.

Finely Tuned Language Models for Enhanced Translation

0

Translation plays an essential role in enabling companies to expand across borders, with requirements varying significantly in terms of tone, accuracy, and technical terminology handling. The emergence of sovereign AI has highlighted critical challenges in large language models (LLMs), particularly their struggle to capture nuanced cultural and linguistic contexts beyond English-dominant frameworks. As global communication becomes increasingly complex, organizations must carefully evaluate translation solutions that balance technological efficiency with cultural sensitivity and linguistic precision.

In this post, we explore how LLMs can address the following two distinct English to Traditional Chinese translation use cases:

  • Marketing content for websites: Translating technical text with precision while maintaining a natural promotional tone.
  • Online training courses: Translating slide text and markdown content used in platforms like Jupyter Notebooks, ensuring accurate technical translation and proper markdown formatting such as headings, sections, and hyperlinks.

These use cases require a specialized approach beyond general translation. While prompt engineering with instruction-tuned LLMs can handle certain contexts, more refined tasks like these often do not meet expectations. This is where fine-tuning Low-Rank Adaptation (LoRA) adapters separately on collected datasets specific to each translation context becomes essential.

Implementing LoRA adapters for domain-specific translation

For this project, we are using Llama 3.1 8B Instruct as the pretrained model and implementing two models fine-tuned with LoRA adapters using NVIDIA NeMo Framework. These adapters were trained on domain-specific datasets—one for marketing website content and one for online training courses. For easy deployment of LLMs with simultaneous use of multiple LoRA adapters on the same pretrained model, we are using NVIDIA NIM.

Refer to the Jupyter Notebook to guide you through executing LoRA fine-tuning with NeMo.

Optimizing LLM deployment with LoRA and NVIDIA NIM

NVIDIA NIM introduces a new level of performance, reliability, agility, and control for deploying professional LLM services. With prebuilt containers and optimized model engines tailored for different GPU types, you can easily deploy LLMs while boosting service performance. In addition to popular pretrained models including the Meta Llama 3 family and Mistral AI Mistral and Mixtral models, you can integrate and fine-tune your own models with NIM, further enhancing its capabilities.

LoRA is a powerful customization technique that enables efficient fine-tuning by adjusting only a subset of the model’s parameters. This significantly reduces required computational resources. LoRA has become popular due to its effectiveness and efficiency. Unlike full-parameter fine-tuning, LoRA adapter weights are smaller and can be stored separately from the pretrained model, providing greater flexibility in deployment.

NVIDIA TensorRT-LLM has established a mechanism that can simultaneously serve multiple LoRA adapters on the same pretrained model. This multi-adapter mechanism is also supported by NIM.

Step-by-step LoRA fine-tuning deployment with NVIDIA LLM NIM

This section describes the three steps involved in LoRA fine-tuning deployment using NVIDIA LLM NIM.

Step 1: Set up the NIM instance and LoRA models

First, launch a computational instance equipped with two NVIDIA L40S GPUs as recommended in the NIM support matrix.

Next, upload the two fine-tuned NeMo files to this environment. Detailed examples of LoRA fine-tuning using NeMo Framework are available in the official documentation and a Jupyter Notebook.

To organize the environment, use the following command to create directories for storing the LoRA adapters:

$ mkdir -p loras/llama-3.1-8b-translate-course
$ mkdir -p loras/llama-3.1-8b-translate-web
$ export LOCAL_PEFT_DIRECTORY=$(pwd)/loras
$ chmod -R 777 $(pwd)/loras
$ tree loras
loras
├── llama-3.1-8b-translate-course
│ └── course.nemo
└── llama-3.1-8b-translate-web
└── web.nemo

2 directories, 2 files

Step 2: Deploy NIM and LoRA models

Now, you can proceed to deploy the NIM container. Replace with your actual NGC API token. Generate an API key if needed. Then run the following commands:

$ export NGC_API_KEY=
$ export LOCAL_PEFT_DIRECTORY=$(pwd)/loras
$ export NIM_PEFT_SOURCE=/home/nvs/loras
$ export CONTAINER_NAME=nim-llama-3.1-8b-instruct

$ export NIM_CACHE_PATH=$(pwd)/nim-cache
$ mkdir -p “$NIM_CACHE_PATH”
$ chmod -R 777 $NIM_CACHE_PATH

$ echo “$NGC_API_KEY” | docker login nvcr.io –username ‘$oauthtoken’ –password-stdin
$ docker run -it –rm –name=$CONTAINER_NAME \
–runtime=nvidia \
–gpus all \
–shm-size=16GB \
-e NGC_API_KEY=$NGC_API_KEY \
-e NIM_PEFT_SOURCE \
-v $NIM_CACHE_PATH:/opt/nim/.cache \
-v $LOCAL_PEFT_DIRECTORY:$NIM_PEFT_SOURCE \
-p 8000:8000 \
nvcr.io/nim/meta/llama-3.1-8b-instruct:1.1.2

After executing these steps, NIM will load the model. Once complete, you can check the health status and retrieve the model names for both the pretrained model and LoRA models using the following commands:

# NIM health status
$ curl http://:8000/v1/health/re…uned-645×399.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/02/bleu-scores-test-datasets-base-model-lora-fine-tuned-485×300.png 485w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/02/bleu-scores-test-datasets-base-model-lora-fine-tuned-146×90.png 146w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/02/bleu-scores-test-datasets-base-model-lora-fine-tuned-362×224.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/02/bleu-scores-test-datasets-base-model-lora-fine-tuned-178×110.png 178w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/02/bleu-scores-test-datasets-base-model-lora-fine-tuned-1024×633.png 1024w” sizes=”(max-width: 1200px) 100vw, 1200px”/>Figure 1. BLEU scores (higher is better) of different test datasets using the base model and two LoRA fine-tuned models

A graph showing COMET scores for the Course and Web dataset, comparing base model with LoRA.
1...323324325...741Page 324 of 741