Home Blog Page 504

Fiverr’s AI Tool Boosts Freelancing Chances

0

Artificial Intelligence Recruiting Tools: Can They Help Freelancers Navigate a Competitive Job Market?

Fiverr’s Dynamic Matching: A Game-Changer in Freelance Recruitment?

Artificial intelligence (AI) recruiting tools keep coming, but can they help freelancers navigate a competitive job market? On Wednesday, the freelancing platform Fiverr released Dynamic Matching, an AI-enhanced tool that it says "transforms how businesses can find and hire freelancers."

How It Works

Businesses can upload an existing project brief or provide a few lines that Dynamic Matching can expand into a full description. After finalizing the brief, the feature scans for applicable freelancers on Fiverr’s platform, and "within about an hour, qualified freelancers will send personalized proposals."

Fiverr’s AI-Driven Approach

In an exclusive interview with ZDNET, Fiverr’s VP of Product Adi Margolis explained that the tool’s oversight team continually optimizes the algorithms that drive it using current data and feedback to ensure it’s "increasingly reliable and precise." Fiverr can customize the feature for Pro users and other higher-tier customers with specific needs.

Addressing Bias Concerns

When asked about any efforts Fiverr makes to ensure that freelancers aren’t disadvantaged by information unrelated to their professional skills, Margolis said, "Our AI model knows how to take only relevant information from the freelancers’ data like their skill set and details from past orders. The models don’t add new information that isn’t accurate."

Current Status and Future Launch

Dynamic Matching is currently in beta and available to Fiverr Pro users. It is set to launch fully early next year.

Conclusion

Fiverr’s Dynamic Matching is an innovative AI-driven tool that has the potential to revolutionize the way businesses find and hire freelancers. By combining AI’s efficiency and data-driven insights with human expertise, the feature ensures that businesses receive the most relevant matches. As the freelance market continues to evolve, it will be interesting to see how AI recruiting tools like Dynamic Matching shape the future of freelance recruitment.

FAQs

Q: What is Fiverr’s Dynamic Matching?
A: Fiverr’s Dynamic Matching is an AI-enhanced tool that "transforms how businesses can find and hire freelancers."

Q: How does Dynamic Matching work?
A: Businesses can upload an existing project brief or provide a few lines that Dynamic Matching can expand into a full description. The feature then scans for applicable freelancers on Fiverr’s platform and provides qualified freelancers with personalized proposals.

Q: Is Dynamic Matching available to all Fiverr users?
A: No, Dynamic Matching is currently available to Fiverr Pro users only. It is set to launch fully early next year.

Q: Does Fiverr’s AI model address bias concerns?
A: Yes, Fiverr’s AI model is designed to only take relevant information from freelancers’ data and does not add new information that isn’t accurate.

Creating a John Wick-inspired Mocap for SPINE: When Unreal Engine 5 Isn’t Always Best

0

Creating Animation for Video Games

Creating animation for video games has become reliant on motion capture, but SPINE developer Nekki has gone further than most, having created its own app, called Cascadeur. This adds to a workflow that includes Unreal Engine 5 and Epic Game’s MetaHuman, but one that doesn’t always rely on those ‘off the shelf’ apps for its stylized anime-like animation.

SPINE’s Lead Animator Talks Performance Capture

In our developer series on upcoming gun-fu game SPINE, we’ve already discovered how Nekki is creating a new ‘intelligent’ action camera as well as why the game’s art director loves using Unreal Engine 5, but here I chat to the game’s lead animator, Evgeniy Khapugin, and discover how performance capture is being used for the game’s fluid, athletic combat.

Challenges of Integrating Motion Capture

The biggest challenge for us is the leap in animation quality, which demands more complex and polished rigs. For comparison, in our previous projects, we had strict limits on the number of joints. In our mobile project Shadow Fight 4: Arena, we used 91 joints, 58 of which were facial. Now, developing for PC and consoles, we no longer have such hardware limitations and have increased the number of joints significantly. Our main character currently has 1,217 joints, 839 of which are facial.

Tools and Approaches for Mocap Data

We use every motion capture tool we can get our hands on! Mocap is a convenient tool that allows us to quickly see results. By the final stages, there may be little to none of the original mocap left, but it lays a strong foundation. Mocap also helps us determine whether something needs to be changed or completely redone.

We also work with external mocap studios using classic optical marker systems, specifically Vicon in our case. There, we capture movements that our team can’t perform – complex falls with rolls, parkour and stunt work, fight choreography, paired combat interactions and interactions involving props.

Unbaking and Animation Refining

Cascadeur has evolved into a standalone product, driven by both internal and external user requests for mocap tools. One such tool is ‘animation unbaking’, which simplifies the process of working with motion capture. This tool functions somewhat similarly to the Simplify Curve filter in Maya, but it’s more advanced.

With internal algorithms, Cascadeur identifies keyframes and retains only those, preserving the overall motion while reducing the number of frames. This makes it easier for animators to edit. After ‘unbaking’, we can work on retiming to make movements sharper. Then, we can apply auto-physics to correct any inaccuracies, and finally, refine amplitudes, arcs, action lines, and so on.

Balance Between Raw Data and Style

In SPINE, we have two types of animation tasks: gameplay animations and cutscene / acting animations. Let’s start with gameplay animations. In gameplay, the most important aspect is how the animation feels and plays in the game, and only then how it looks. If you view these animations outside of the game, they may not look as physically accurate as if they were in a cutscene.

When the player presses a button, they expect an immediate response, and the animation needs to provide that feedback. Raw motion capture data is useful in the early stages, just adjusting the timing and distances. We insert it roughly, without fine-tuning. The focus here is speed and iteration. Playing these rough animations with the game designer helps us figure out the direction.

Unreal Engine 5 and Procedural Animations

We don’t use Unreal Engine 5 during the mocap process. When working with Vicon, we use MotionBuilder to view the character immediately, while for Xsens, we rely on their native software without retargeting. In our animation pipeline, we try to avoid using Unreal for animations. Not because Unreal is bad, but because it’s important for us to have a single source of animation software.

In Unreal Engine 5, we handle procedural animations. For example, our robotic spiders walk procedurally. They have six legs, climb walls and floors, and their legs automatically find where to step and react accordingly.

Conclusion

Nekki’s use of motion capture and custom software, such as Cascadeur, has allowed for a high level of detail and realism in the animation process for SPINE. By combining these tools with Unreal Engine 5, the team has been able to create a unique and engaging gaming experience.

FAQs

Q: How does Nekki approach the animation process for SPINE?

A: Nekki uses a combination of motion capture, custom software, and Unreal Engine 5 to create high-quality animations for SPINE. The team focuses on capturing realistic movements and then stylizes them to fit the game’s aesthetic and gameplay requirements.

Q: What is Cascadeur and how does it aid the animation process?

A: Cascadeur is a custom software developed by Nekki for motion capture and animation processing. It allows for animation unbaking, retiming, and auto-physics, making it easier to refine and adjust animations.

Q: How does Nekki balance the need for realism and stylization in animation?

A: Nekki approaches animation as a two-step process, first capturing realistic movements with motion capture and then stylizing them to fit the game’s aesthetic and gameplay requirements. The team also adjusts timing and distances to make the animations feel more dynamic and engaging.

Accelerate Container Security with NVIDIA NIM Agent Blueprint

0

Addressing Software Security Issues with Generative AI

Challenges in Software Security

Addressing software security issues is becoming increasingly challenging as the number of vulnerabilities reported in the CVE database continues to grow at an accelerated pace. With over 200,000 vulnerabilities reported at the end of 2023, the traditional approach to scanning and patching has become unmanageable. Assessing a single container for vulnerabilities requires the collection, comprehension, and synthesis of hundreds of pieces of information.

The Rise of Generative AI

Enterprises are increasingly adopting generative AI to drive innovation across domains. Vulnerability detection and resolution will become a top generative AI use case in software delivery, according to the IDC. Generative AI can improve vulnerability defense while reducing the burden on security teams. Organizations have already begun to explore its use for automation, but scaling it at an enterprise level requires a complex AI system.

Accelerating Vulnerability Analysis with Generative AI

Video 1 shows how NVIDIA uses generative AI and retrieval-augmented generation (RAG) to accelerate vulnerability analysis in software containers at enterprise scale, dramatically reducing the time to assess and mitigate CVEs from hours or days to mere seconds.

Key Takeaways

  • Using NVIDIA NIM and the NVIDIA Morpheus cybersecurity AI SDK, this event-driven RAG example can dramatically decrease CVE analysis and remediation from days to just seconds.
  • LLM agents can expedite investigations and cut through the noise of an increasing number of known CVEs to highlight urgent security risks.
  • Implementing multiple LLM agents can automate vulnerability management, verification, and VEX justification, all triggered by the results of upstream vulnerability scans.
  • The NIM Agent Blueprint uses asynchronous and parallel GPU processing for scalable, fast analysis of multiple CVEs simultaneously, enabling real-time insights into container and vulnerability information, streamlining the validation process and addressing potential security threats.

Summary

Try the blueprint for free at build.nvidia.com. Learn more and get notified of the upcoming release of a downloadable vulnerability analysis NIM Agent blueprint.

Frequently Asked Questions

Q: What is the purpose of the NVIDIA Morpheus cybersecurity AI SDK?
A: The NVIDIA Morpheus cybersecurity AI SDK is designed to provide a framework for building AI-powered cybersecurity solutions.

Q: How does the NIM Agent Blueprint use generative AI?
A: The NIM Agent Blueprint uses retrieval-augmented generation (RAG) to accelerate vulnerability analysis in software containers at enterprise scale.

Q: What is the benefit of using LLM agents in vulnerability detection?
A: LLM agents can expedite investigations and cut through the noise of an increasing number of known CVEs to highlight urgent security risks.

Apple Completes AI Starter Kit

0

Apple Intelligence: A Mixed Bag of AI Features

Suggested replies aren’t new in iOS 18.2, but they’re a piece of the Apple Intelligence feature set that’s falling into place with this week’s public release of 18.2. Those suggestions I got while planning lunch kind of sum up my whole experience with Apple’s AI up ’til now: occasionally helpful, sometimes way off base, and often good for a laugh. But once the novelty wears off, it’s easily ignored — just like the AI feature sets on every other so-called AI smartphone I’ve used this year.

Apple’s AI Journey

Apple took its time getting here. The first set of AI features dropped with iOS 18.1 at the end of October, including notification and email summaries, generative writing tools, and a cleanup tool to take distractions out of photos. It felt like a deeply minimum viable product, but Apple had to get something out the door for its “built for Apple Intelligence” iPhones.

iOS 18.2: A Meatier Set of Updates

Now, iOS 18.2 has officially arrived after months of beta testing with a meatier set of updates: the Image Playground app for AI image generation, Genmoji, and a ChatGPT extension for Siri. You also get Visual Intelligence, but only with an iPhone 16 or 16 Pro, for reasons that are unclear. There’s more to come, of course, but Apple has finally shipped a set of AI features that resembles Samsung’s and Google’s. The problem is that all of those phone makers are still a long way from delivering the AI smartphones we’ve been promised.

Siri’s Big Update

Siri’s big update in 18.2 is the addition of ChatGPT. It’ll still set timers and answer your basic questions the way it always has, but now it can send more complex queries to ChatGPT. It’s opt in and doesn’t require an OpenAI account to use, which is nice. It’s still just as prone to making stuff up as ever, but it can act as a helpful starting point if you want some assistance with a complex topic.

Image Playground: A Flashy but Limited Feature

Of all the updates iOS 18.2 offers, Image Playground is probably the flashiest. It’s a standalone app with a waitlist, but once you’re in, it unlocks image creation tools in other places throughout the OS, too. Image Playground is a lot like Google’s Pixel Studio, but with way stricter guardrails — that’s a good thing, mostly. Requests to create an image of Pikachu sticking a paper clip in an electrical outlet were denied, which is great news for Pikachu.

Genmoji: A Feature with Limitations

Genmoji is on even stricter rails, and in my experience, it gets things right a lot. But there are some pretty obvious limitations, including the fact that they’re so tiny it’s hard to see much detail in them. They’re supposed to be tiny, but you can forget that and get carried away adding a bunch of stuff and then find it all barely visible in the final product. It made a decent depiction of me in front of a Christmas tree drinking from a red coffee cup that looks good in preview but is impossible to parse at typical emoji size. They also don’t work well in group texts with RCS, so I can’t respond to family texts with obnoxious emoji, which is my primary use case for this feature.

Conclusion

That’s my biggest problem with AI on phones right now. Often, it does what it’s supposed to do. But it’s rarely helpful and doesn’t feel like it’s solving any real problem I was having. That’s been my complaint about this year’s devices from Google and Samsung; now, Apple is at least in the conversation. But they’re all in the same position, with equal pressure to deliver something in 2025 that isn’t just a collection of funny tricks — the novelty is wearing off fast.

FAQs

Q: What is Apple Intelligence?
A: Apple Intelligence is a set of AI features that Apple has been developing for its iPhones, including suggested replies, image generation, and more.

Q: What is Image Playground?
A: Image Playground is a standalone app that allows users to create images using AI technology. It’s a part of the Apple Intelligence feature set.

Q: What is Genmoji?
A: Genmoji is a feature that allows users to create custom emojis using AI technology. It’s a part of the Apple Intelligence feature set.

Q: Is Siri getting an update in iOS 18.2?
A: Yes, Siri is getting an update in iOS 18.2, which includes the addition of ChatGPT. It’ll allow users to send more complex queries to ChatGPT and get more accurate answers.

ChatGPT’s Advanced Voice Mode Gets Visual Context

0

What are the ’12 days of OpenAI’?

OpenAI CEO Sam Altman shared a bit more details about the event, which kicked off at 10 a.m. PT on Dec. 5 and will occur daily for 12 weekdays with a live stream featuring a launch or demo. The launches will be both “big ones” or “stocking stuffers,” according to Altman.

What has been dropped so far?

Thursday, December 12

OpenAI addressed the elephant in the room — the fact that the company’s live stream went down the day before. OpenAI apologized for the inconvenience and said its team is working on a post-mortem to be posted later today.

  • Advanced Voice Mode now has screen-sharing and visual capabilities, meaning it can assist with the context of what it is viewing, whether that be from your phone camera or what’s on your screen.
  • These capabilities build on what Advanced Voice could already do very well — engaging in casual conversation as a human would. The natural-like conversations can be interrupted, have multi-turns, and understand non-linear trains of thought.
  • In the demo, the user gets directions from ChatGPT’s Advanced Voice on how to make a cup of coffee. As the demoer goes through the steps, ChatGPT is verbally offering insights and directions.
  • There’s another bonus for the Christmas season: Users can access a new Santa voice. To activate it, all users have to do is click on the snowflake icon. Santa is rolling out throughout today everywhere that users can access ChatGPT voice mode. The first time you talk to Santa, your usage limits reset, even if you have reached the limit already, so you can have a conversation with him.
  • Video and screen sharing are rolling out in the latest mobile apps starting today and throughout next week to all Team users and most Pro and Plus subscribers. Pro and Plus subscribers in Europe will get access “as soon as we can,” and Enterprise and Edu users will get access early next year.

Wednesday, December 11

Apple released iOS 18.2 today. The release includes integrations with ChatGPT across Siri, Writing Tools, and Visual Intelligence. As a result, today’s live stream focused on walking through the integration.

  • Siri can now recognize when you ask questions outside its scope that could benefit from being answered by ChatGPT instead. In those instances, it will ask if you’d like to process the query using ChatGPT. Before any request is sent to ChatGPT, a message notifying the user and asking for permission will always appear, placing control in the user’s hands as much as possible.
  • Visual Intelligence refers to a new feature for the iPhone 16 lineup that users can access by tapping the Camera Control button. Once the camera is open, users can point it at something and search the web with Google, or use ChatGPT to learn more about what they are viewing or perform other tasks such as translating or summarizing text.
  • Writing Tools now features a new “Compose” tool, which allows users to create text from scratch by leveraging ChatGPT. With the feature, users can even generate images using DALL-E.

All of the above features are subject to ChatGPT’s daily usage limits, the same way that users would reach limits while using the free version of the model on ChatGPT. Users can choose whether or not to enable the ChatGPT integration in Settings.

Tuesday, December 10

  • Canvas is coming to all web users, regardless of plan, in GPT-4o, meaning it is no longer just available in beta for ChatGPT Plus users.
  • Canvas has been built into GPT-4o natively, meaning you can just call on Canvas instead of having to go to the toggle on the model selector.
  • The Canvas interface is the same as what users saw in beta in ChatGPT Plus, with a table on the left-hand side that shows the Q+A exchange and a right-hand tab that shows your project, displaying all of the edits as they go, as well as shortcuts.
  • Canvas can also be used with custom GPTs. It is turned on by default when creating a new one, and there is an option to add Canvas to existing GPTs.
  • Canvas also has the ability to run Python code directly in Canvas, allowing ChatGPT to execute coding tasks such as fixing bugs.

Monday, December 9

OpenAI teased the third-day announcement as “something you’ve been waiting for,” followed by the much-anticipated drop of its video model — Sora.

  • Known as Sora Turbo, the video model is smarter than the February model that was previewed.
  • Access is coming in the US later today; users need only ChatGPT Plus and Pro.
  • Sora can generate video-to-video, text-to-video, and more.
  • ChatGPT Plus users can generate up to 50 videos per month at 480p resolution or fewer videos at 720p. The Pro Plan offers 10x more usage.
  • The new model is smarter and cheaper than the previewed February model.
  • Sora features an explore page where users can view each other’s creations. Users can click on any video to see how it was created.
  • A live demo showed the model in use. The demo-ers entered a prompt and picked aspect ratio, duration, and even presets. I found the live demo video results to be realistic and stunning.
  • OpenAI also unveiled Storyboard, a tool that lets users generate inputs for every frame in a sequence.

Friday, December 6

On the second day of “shipmas,” OpenAI expanded access to its Reinforcement Fine-Tuning Research Program:

  • The Reinforcement Fine-Tuning program allows developers and machine learning engineers to fine-tune OpenAI models to “excel at specific sets of complex, domain-specific tasks,” according to OpenAI.
  • Reinforcement Fine-Tuning refers to a customization technique in which developers can define a model’s behavior by inputting tasks and grading the output. The model then uses this feedback as a guide to improve, becoming better at reasoning through similar problems, and enhancing overall accuracy.
  • OpenAI encourages research institutes, universities, and enterprises to apply to the program, particularly those that perform narrow sets of complex tasks, could benefit from the assistance of AI, and perform tasks that have an objectively correct answer.
  • Spots are limited; interested applicants can apply by filling out this form.
  • OpenAI aims to make Reinforcement Fine-Tuning publicly available in early 2025.

Thursday, December 5

OpenAI started with a bang, unveiling two major upgrades to its chatbot: a new tier of ChatGPT subscription, ChatGPT Pro, and the full version of the company’s o1 model.

The full version of o1:

  • Will be better for all kinds of prompts, beyond math and science
  • Will make major mistakes about 34% less often than o1-preview, while thinking about 50% faster
  • Rolls out today, replacing o1-preview to all ChatGPT Plus and now Pro users
  • Lets users input images, as seen in the demo, to provide multi-modal reasoning (reasoning on both text and images)

ChatGPT Pro:

  • Is meant for ChatGPT Plus superusers, granting them unlimited access to the best OpenAI has to offer, including unlimited access to OpenAI o1-mini, GPT-4o, and Advanced Mode
  • Features o1 pro mode, which uses more computing to reason through the hardest science and math problems
  • Costs $200 per month

Where can you access the live stream?

The live streams are held on the OpenAI website, and posted to its YouTube channel immediately after. To make access easier, OpenAI will also post a link to the live stream on its X account 10 minutes before it starts, which will be at approximately 10 a.m. PT/1 p.m. ET daily.

What can you expect?

The releases remain a surprise, but many anticipate that Sora, OpenAI’s video model initially announced last February, will be launched as part of one of the bigger drops. Since that first announcement, the model has been available to a select group of red teamers and testers and was leaked last week by some testers over grievances about “unpaid labor,” according to reports.


As OpenAI’s “12 Days of OpenAI” continue, it’s clear that the company is dedicated to pushing the boundaries of what’s possible with AI. From advancements in ChatGPT to the introduction of new features like Advanced Voice Mode and Canvas, the event has been filled with exciting announcements. As we head into the final stretch

Character.AI has Retrained its Chatbots to Stop Chatting up Teens

0

Character.AI Announces Parental Controls for Teen Users

In an announcement today, Chatbot service Character.AI says it will soon be launching parental controls for teenage users, and it described safety measures it’s taken in the past few months, including a separate large language model (LLM) for users under 18. The announcement comes after press scrutiny and two lawsuits that claim it contributed to self-harm and suicide.

New Safety Features

In a press release, Character.AI said that, over the past month, it’s developed two separate versions of its model: one for adults and one for teens. The teen LLM is designed to place “more conservative” limits on how bots can respond, “particularly when it comes to romantic content.” This includes more aggressively blocking output that could be “sensitive or suggestive,” but also attempting to better detect and block user prompts that are meant to elicit inappropriate content. If the system detects “language referencing suicide or self-harm,” a pop-up will direct users to the National Suicide Prevention Lifeline, a change that was previously reported by The New York Times.

Minors’ Interactions Limited

Minors will also be prevented from editing bots’ responses — an option that lets users rewrite conversations to add content Character.AI might otherwise block.

Additional Features

Beyond these changes, Character.AI says it’s “in the process” of adding features that address concerns about addiction and confusion over whether the bots are human, complaints made in the lawsuits. A notification will appear when users have spent an hour-long session with the bots, and an old disclaimer that “everything characters say is made up” is being replaced with more detailed language. For bots that include descriptions like “therapist” or “doctor,” an additional note will warn that they can’t offer professional advice.

Parental Control Options

The parental control options are coming in the first quarter of next year, Character.AI says, and they’ll tell parents how much time a child is spending on Character.AI and which bots they interact with most frequently. All the changes are being made in collaboration with “several teen online safety experts,” including the organization ConnectSafely.

Conclusion

Character.AI is taking steps to address concerns about the safety and well-being of its teenage users. The new safety features and parental control options aim to provide a safer and more responsible experience for minors. While the company still faces lawsuits and criticism, these changes demonstrate its commitment to continuously improving its policies and product.

FAQs

Q: What are the new safety features?
A: The new safety features include a separate large language model (LLM) for users under 18, more aggressive blocking of sensitive or suggestive content, and better detection and blocking of user prompts meant to elicit inappropriate content.

Q: How will minors be prevented from editing bots’ responses?
A: Minors will be prevented from editing bots’ responses to prevent them from adding content that Character.AI might otherwise block.

Q: What are the parental control options?
A: The parental control options will allow parents to track how much time their child is spending on Character.AI and which bots they interact with most frequently.

Q: When will the parental control options be available?
A: The parental control options will be available in the first quarter of next year.

Q: Who is collaborating with Character.AI on these changes?
A: Character.AI is collaborating with several teen online safety experts, including the organization ConnectSafely.

Momentum Matters

0

Momentum in Gradient Descent

We often think of Momentum as a means of dampening oscillations and speeding up the iterations, leading to faster convergence. But it has other interesting behavior.

Step-size α = 0.02

Momentum β = 0.99

We often think of Momentum as a means of dampening oscillations and speeding up the iterations, leading to faster convergence. But it has other interesting behavior. It allows a larger range of step-sizes to be used, and creates its own oscillations. What is going on?

Here’s a popular story about momentum: gradient descent is a man walking down a hill. He follows the steepest path downwards; his progress is slow, but steady. Momentum is a heavy ball rolling down the same hill. The added inertia acts both as a smoother and an accelerator, dampening oscillations and causing us to barrel through narrow valleys, small humps and local minima.

This standard story isn’t wrong, but it fails to explain many important behaviors of momentum. In fact, momentum can be understood far more precisely if we study it on the right model.

One nice model is the convex quadratic. This model is rich enough to reproduce momentum’s local dynamics in real problems, and yet simple enough to be understood in closed form. This balance gives us powerful traction for understanding this algorithm.

Gradient Descent

Gradient descent has many virtues, but speed is not one of them. It is simple — when optimizing a smooth function f, we make a small step in the gradient:

w^k+1 = w^k – α ∇f(w^k).

For a step-size small enough, gradient descent makes a monotonic improvement at every iteration. It always converges, albeit to a local minimum. And under a few weak curvature conditions it can even get there at an exponential rate.

Momentum

But the exponential decrease, though appealing in theory, can often be infuriatingly small. Things often begin quite well — with an impressive, almost immediate decrease in the loss. But as the iterations progress, things start to slow down. You start to get a nagging feeling you’re not making as much progress as you should be. What has gone wrong?

The problem could be the optimizer’s old nemesis, pathological curvature. Pathological curvature is, simply put, regions of f which aren’t scaled properly. The landscapes are often described as valleys, trenches, canals and ravines. The iterates either jump between valleys, or approach the optimum in small, timid steps. Progress along certain directions grind to a halt. In these unfortunate regions, gradient descent fumbles.

Momentum proposes the following tweak to gradient descent. We give gradient descent a short-term memory:

z^k+1 = βz^k + ∇f(w^k).

Conclusion

Momentum is often misunderstood as simply a way to speed up gradient descent. But it has much deeper implications for the behavior of the algorithm. By introducing a short-term memory, momentum can help the algorithm navigate pathological curvature and achieve faster convergence.

FAQs

What is momentum in gradient descent?
Momentum is a modification to the gradient descent algorithm that introduces a short-term memory. It helps the algorithm navigate pathological curvature and achieve faster convergence.

How does momentum work?
Momentum works by adding a component to the update step that is proportional to the previous update step. This helps the algorithm to build up momentum and navigate the optimization landscape more effectively.

What are the benefits of momentum?
The benefits of momentum include faster convergence, improved stability, and better performance on optimization problems with non-smooth or non-convex objectives.

Are there any drawbacks to using momentum?
Yes, one drawback to using momentum is that it can make the algorithm more sensitive to the choice of hyperparameters. Additionally, momentum can sometimes get stuck in local minima.

How do I choose the right step-size and momentum parameters?
The choice of step-size and momentum parameters depends on the specific optimization problem and the desired performance of the algorithm. It is often a good idea to use cross-validation or other methods to select the best parameters.

Can I use momentum with other optimization algorithms?
Yes, momentum can be used with other optimization algorithms, such as stochastic gradient descent or Adam. However, the specific implementation and hyperparameters may need to be adjusted depending on the algorithm being used.

Critical WordPress Plugin Vulnerability Under Active Exploit

0

Thousands of WordPress Sites Remain Unpatched Against Critical Security Flaw

Significant, Multifaceted Threat

Thousands of sites running WordPress remain unpatched against a critical security flaw in a widely used plugin that was being actively exploited in attacks that allow for unauthenticated execution of malicious code, security researchers said.

The Vulnerability

The vulnerability, tracked as CVE-2024-11972, is found in Hunk Companion, a plugin that runs on 10,000 sites that use the WordPress content management system. The vulnerability, which carries a severity rating of 9.8 out of a possible 10, was patched earlier this week. At the time this post went live on Ars, figures provided on the Hunk Companion page indicated that less than 12 percent of users had installed the patch, meaning nearly 9,000 sites could be next to be targeted.

A Serious Concern for Site Integrity

"This vulnerability represents a significant and multifaceted threat, targeting sites that use both a ThemeHunk theme and the Hunk Companion plugin," Daniel Rodriguez, a researcher with WordPress security firm WP Scan, wrote. "With over 10,000 active installations, this exposed thousands of websites to anonymous, unauthenticated attacks capable of severely compromising their integrity."

The Exploit

WP Scan discovered the vulnerability while analyzing the compromise of a customer’s site. The firm found that the initial vector was CVE-2024-11972. The exploit allowed the hackers behind the attack to cause vulnerable sites to automatically navigate to wordpress.org and download WP Query Console, a plugin that hasn’t been updated in years.

Conclusion

The discovery of this vulnerability highlights the importance of timely patching and regular security audits for WordPress websites. With the sheer number of unpatched sites, it is crucial that site administrators take immediate action to prevent potential attacks and compromise.

Frequently Asked Questions

Q: What is the vulnerability?
A: The vulnerability is a critical security flaw in the Hunk Companion plugin, tracked as CVE-2024-11972, which allows for unauthenticated execution of malicious code.

Q: How many sites are affected?
A: The vulnerability affects over 10,000 sites that use the Hunk Companion plugin.

Q: Is the patch available?
A: Yes, the patch was released earlier this week, and site administrators are urged to install it as soon as possible.

Q: What should I do if I’m affected?
A: Install the patch immediately, and consider conducting a security audit to identify any potential vulnerabilities.

Sora’s AI Video Revolution

0

The First Version of OpenAI’s Sora: A Promising but Flawed AI Video Generator

Sora’s Capabilities and Limitations

OpenAI’s Sora, a new AI video generator, has been released after almost a year of teasers. The platform can generate videos of just about anything, from superheroes to cityscapes to animated puppies. However, the results are far from satisfactory, with many videos plagued by oddities and inconsistencies.

Getting Started with Sora

To access Sora’s features, users must create an account, which was closed due to overwhelming demand. A $20 monthly "Plus" membership is required to generate videos at 480p or 720p, capped at either five or 10 seconds in length. To unlock everything, including 1080p quality and 20-second-long videos, users must pay $200 a month for the "Pro" subscription.

Testing Sora’s Video Generation

My results from testing the Plus tier have been underwhelming. Simple prompts with limited descriptions seem to work best, such as "a cat playing with a ball of yarn." However, Sora often adds unwanted elements, like a second tail for a few moments, or inserts CGI that looks jittery and unnatural.

Challenges with Complex Prompts

Complex prompts with detailed scene descriptions are even more problematic. It’s difficult to achieve natural human motion, with hands flailing everywhere when asking Sora to show someone applying makeup, and videos of people eating salad and sausage rolls resembling viral AI clips of Will Smith inhaling spaghetti.

The Storyboard Feature

Sora’s Storyboard feature allows users to explain what they want Sora to generate every two seconds, similar to a video editing timeline. While easy to use, the results are still poor, with more distortions and weirdness appearing the more detail added.

Some Impressions and Limitations

Some things do impress, such as video generation speed, which is generally under 30 seconds for even 10-second-long clips. Patterns on fur and textiles remain consistent, even during fast-paced movement, and lighting, shadow, and mirror effects simulate real-life conditions. Sunlight coming through a window produces a flash of glare and shines through materials as expected. However, most objects have high levels of detail and don’t pixelate.

Comparison with Runway AI

Sora outperforms Runway AI, considered one of the better AI video generators for simulating photorealism. When using identical prompts, Sora’s results look more realistic and have fewer visual distortions.

Conclusion

While Sora shows promise, it’s far from ready for entertainment or commercial work that requires narrative coherence. Even experienced users may struggle to produce high-quality videos that don’t include obvious AI weirdness. The platform’s current limitations and high subscription costs make it inaccessible for many.

FAQs

Q: Is Sora suitable for entertainment or commercial work?
A: No, it’s not ready for narrative coherence.

Q: Can I use Sora to create high-quality videos?
A: No, the results are heavily plagued by oddities and inconsistencies.

Q: Is Sora accessible for everyone?
A: No, the high subscription costs make it inaccessible for many.

Q: Can I use Sora to create content for young children?
A: Yes, but be aware that Sora can generate nonsensical AI-generated content targeted towards young children.

Mistral-NeMo-Minitron 8B: Unparalleled Accuracy

0

Mistral NeMo 12B: A State-of-the-Art Large Language Model

Introduction

Recently, NVIDIA and Mistral AI unveiled Mistral NeMo 12B, a leading state-of-the-art large language model (LLM). Consistently outperforming similarly sized models on a wide range of benchmarks, Mistral NeMo 12B has set a new standard for language modeling.

Mistral NeMo Minitron 8B

Building on the success of Mistral NeMo 12B, we announced Mistral-NeMo-Minitron 8B, one of the most advanced open-access models in its size class. This model consistently delivers leading accuracy on nine popular benchmarks. The Mistral-NeMo-Minitron 8B base model was obtained by width-pruning the Mistral NeMo 12B base model, followed by a light retraining process using knowledge distillation.

Model Pruning and Distillation

Model pruning is the process of making a model smaller and leaner, either by dropping layers (depth pruning) or dropping neurons and attention heads and embedding channels (width pruning). Pruning is often accompanied by some amount of retraining for accuracy recovery. Model distillation is a technique used to transfer knowledge from a large, complex model, often called the teacher model, to a smaller, simpler student model.

Iterative Pruning and Distillation

The combination of model pruning followed by light retraining through distillation is an effective and cost-efficient approach to train a family of models. For each additional model, just 100-400B tokens are used for retraining—a greater than 40x reduction compared to training from scratch.

Best Practices for Structured Weight Pruning and Knowledge Distillation

The learning from extensive ablation studies has been summarized into 10 best practices for structured weight pruning combined with knowledge distillation. We found that width pruning consistently outperforms depth pruning and, most importantly, pruned and distilled models outperform models trained from scratch in quality.

Mistral-NeMo-Minitron 8B

Following our best practices, we width-pruned the Mistral NeMo 12B model to obtain an 8B target model. This section details the steps and parameters used to obtain the Mistral-NeMo-Minitron 8B base model, as well as its performance.

Teacher Fine-Tuning

To correct for the distribution shift across the original dataset, we first fine-tuned the unpruned Mistral NeMo 12B model on our dataset using 127B tokens. Experiments showed that, without correcting for the distribution shift, the teacher provides suboptimal guidance on the dataset when being distilled.

Width-Only Pruning

Given our goal of obtaining the strongest 8B model possible, we proceeded with width-only pruning. We pruned both the embedding (hidden) and MLP intermediate dimensions along the width axis to compress Mistral NeMo 12B.

Distillation Parameters

We distilled the model with peak learning rate=1e-4, minimum learning rate=4.5e-7, linear warm up of 60 steps, cosine decay schedule, and a global batch size of 768 using 380B tokens (the same dataset used in teacher fine-tuning).

Mistral-NeMo-Minitron-8B-Instruct

We applied an advanced alignment technique consisting of two-stage instruction fine-tuning and two-stage preference optimization, resulting in a state-of-the-art instruct model with excellent performance in instruction following, language reasoning, function calling, and safety benchmarks.

Performance Benchmarks

We optimized the Mistral-NeMo-Minitron-8B-Base model, the teacher Mistral-NeMo-12B model, and the LLama-3.1-8B model with NVIDIA TensorRT-LLM, an open-source toolkit for optimized LLM inference.

Conclusion

Mistral-NeMo-Minitron-8B provides class-leading accuracy and consistently outperforms recently introduced state-of-the-art models of similar size. Mistral-NeMo-Minitron-8B is our first work on the distillation of the Mistral-NeMo-12B model and provides strong support for our structured weight pruning combined with knowledge distillation best practices.

Frequently Asked Questions

Q: What is the key innovation in Mistral NeMo 12B?

A: The key innovation is the combination of model pruning followed by light retraining through distillation, which is an effective and cost-efficient approach to train a family of models.

Q: What is the difference between depth pruning and width pruning?

A: Depth pruning involves dropping layers, while width pruning involves dropping neurons and attention heads and embedding channels.

Q: How does Mistral NeMo Minitron 8B compare to other models of similar size?

A: Mistral NeMo Minitron 8B consistently outperforms other models of similar size on a wide range of benchmarks.

Q: How does the performance of Mistral NeMo Minitron 8B change when deployed in FP8 precision?

A: Deployment in FP8 delivers a performance boost of ~1.4x across all three models compared to BF16.