Home Blog Page 617

AI Agent Launch

0

OpenAI Preparing to Release Autonomous AI Agent

OpenAI is preparing to release an autonomous AI agent that can control computers and perform tasks independently, code-named “Operator.” The company plans to debut it as a research preview and developer tool in January, according to Bloomberg.

Competition Heats Up

This move intensifies the competition among tech giants developing AI agents: Anthropic recently introduced its “computer use” capability, while Google is reportedly preparing its own version for a December release. The timing of Operator’s eventual consumer release remains under wraps, but its development signals a pivotal shift toward AI systems that can actively engage with computer interfaces rather than just process text and images.

AI Companies’ Promises

All the leading AI companies have promised autonomous AI agents, and OpenAI has hyped up the possibility recently. In a Reddit “Ask Me Anything” forum a few weeks ago, OpenAI CEO Sam Altman said “we will have better and better models,” but “I think the thing that will feel like the next giant breakthrough will be agents.” At an OpenAI press event ahead of the company’s annual Dev Day last month, chief product officer Kevin Weil said: “I think 2025 is going to be the year that agentic systems finally hit the mainstream.”

Monetizing AI Models

AI labs face mounting pressure to monetize their costly models, especially as incremental improvements may not justify higher prices for users. The hope is that autonomous agents are the next breakthrough product — a ChatGPT-scale innovation that validates the massive investment in AI development.

Conclusion

The release of OpenAI’s autonomous AI agent, Operator, marks a significant milestone in the development of AI systems. As the competition among tech giants heats up, the industry is poised for a breakthrough in AI capabilities. With autonomous agents, AI systems will be able to actively engage with computer interfaces, opening up new possibilities for innovation and growth.

FAQs

Q: What is OpenAI’s autonomous AI agent, Operator?
A: Operator is an autonomous AI agent that can control computers and perform tasks independently.

Q: When will Operator be released?
A: OpenAI plans to debut Operator as a research preview and developer tool in January.

Q: Will Operator be available to consumers?
A: The timing of Operator’s eventual consumer release remains under wraps, but its development signals a pivotal shift toward AI systems that can actively engage with computer interfaces.

Q: What are the implications of autonomous AI agents?
A: Autonomous AI agents will enable AI systems to actively engage with computer interfaces, opening up new possibilities for innovation and growth.

Button Holding: The Superior UX Design Choice

0

Here is the rewritten article:

The Problem with Confirmation Dialogs

“I accidentally deleted… Can you find it in your backups?” says every user ever.

Yeah, “accidentally,” sure…

But since I think that confirmation dialogs are a bad design in general, these requests are annoying, but I get it. It’s just proof that the traditional approach is ineffective in preventing terminal actions a user doesn’t really want to make.

Confirmation pop-ups – whether browser-based or custom – are often more annoying than helpful. They interrupt the user’s flow and can lead to “confirmation fatigue,” where users habitually click through dialogs without truly processing the action they’re about to take.

The Advantage of Hold-to-Confirm Buttons

Frankly, for how simple it is to implement, it’s a game-changer.

Hold-to-confirm buttons require the user to press and hold a button for a predetermined amount of time (e.g., 1 to 5 seconds) to execute an action. This method offers several benefits:

• Intentionality: Holding a button creates a continuous period where the user subconsciously but actively considers their action.
• Reduced Errors: By requiring sustained interaction, accidental clicks are none. Really, zero – we’ve had none since implementation, unlike tens of requests per month with confirmation dialogs.
• Enhanced UX: It provides a smoother, less intrusive experience compared to disruptive confirmation dialogs.

How It Works

When a user initiates a hold-to-confirm action, a visual indicator (like a progress bar) fills up over the hold period. If the user releases the button before completion, the action is canceled. This simple mechanism effectively replaces the need for a secondary confirmation step.

Real-World Implementation and Results

We first implemented hold-to-confirm buttons in our Grace projects for deletion actions. We set a 3-second hold time with a filling progress bar of a contrasting color. This duration proved sufficient without being tedious.

The Outcome?

• Immediate Success: We saw a total mitigation of accidental deletions.
• Positive Feedback: Only a few users needed clarification on the new mechanism.
• Improved Satisfaction: Users appreciated it over dialogs.

Determining the Appropriate Hold Duration

It’s not as tricky as it sounds. Here’s how we go about it:

• Terminal Actions (e.g., Deletions): 3 seconds. These actions are irreversible and may affect user data significantly.
• Critical Actions Affecting Others: Up to 5 seconds. For actions like canceling an invoice that notifies customers, a longer hold time ensures a higher level of deliberate intent. It can get a little bit annoying, which is the purpose.
• Easily Reversible Actions: 1 to 1.5 seconds. For actions like unpublishing articles, shorter hold times maintain efficiency while still requiring intent and thought.

The Importance of Visual Feedback

Visual cues are the most important for the effectiveness of hold-to-confirm buttons:

• Progress Indicators: A filling bar or countdown communicates that the button needs to be held.
• Immediate Feedback: If the button is clicked (not held), a brief movement of the progress bar signals that the user should hold the button.
• Completion Animation: A subtle effect after the hold is complete confirms that the action has been executed.

Implementing Hold-to-Confirm in Your Projects

To implement hold-to-confirm buttons in your projects, follow these steps:

1. Choose a suitable hold duration based on the action’s importance and potential impact.
2. Design a visually appealing progress indicator that communicates the hold requirement.
3. Use immediate feedback to notify the user if they release the button prematurely.
4. Provide a subtle completion animation to confirm the action’s execution.

Conclusion

Hold-to-confirm buttons offer a superior alternative to traditional confirmation dialogs by:

• Enhancing user intentionality.
• Reducing accidental actions.
• Streamlining the user experience.

By rethinking how we approach confirmations, we can create applications that are not only more user-friendly but also more effective in preventing unintended consequences.

FAQs

Q: Have you used hold-to-action buttons?
A: Yes, we’ve implemented hold-to-confirm buttons in our Grace projects and seen a significant reduction in accidental actions.

Q: Do you get those “accidentally deleted” notes from users?
A: No, since implementing hold-to-confirm buttons, we’ve had none of those requests.

Q: What do you think about confirmation dialogs or holding buttons?
A: I think hold-to-confirm buttons are a game-changer. They provide a smoother, more intentional user experience compared to traditional confirmation dialogs.

Google’s Nest Cameras Get AI-Powered Search

0

The Google Smart Home is Getting a Major Upgrade

The Google smart home is getting one of its biggest upgrades in years. In addition to the recent launch of a new Google TV Streamer and Nest Learning Thermostat, the company is bringing Gemini, its set of artificial intelligence (AI) large language models (LLMs), to its smart home experience with new AI-generated summaries, automations, and a Gemini-powered search feature.

Gemini-Powered Nest Surveillance & Search

Beginning today, Google is adding Gemini to the smart home Nest cameras in the Google Home app, using AI-powered image recognition to generate descriptions of what the camera captures. This feature is only available for cameras in a Nest Aware Plus subscription.

Google shared the example of a camera looking into a backyard where kids jump into a swimming pool. The description read, "Three children jumping into a swimming pool, chasing after a pool toy. The sun shines brightly, casting long shadows," instead of an ambiguous "motion detected."

This approach reduces false alerts and provides rich descriptions of captured information. I asked Google if the feature would work with Nest Aware’s person identification tech, which recognizes previously identified people to give you more specific alerts. Google said this option isn’t available yet, but the company is working on it.

Coming Soon: A More Powerful Smart Home Assistant

Nest cameras’ Gemini-powered alerts will let you create automations based on what is detected in the image. For example, you can ask Google Assistant if your kids left their bikes on the driveway and then tell Assistant to remind them to bring their bikes into the garage when they arrive from school. Gemini will set an automation through a camera to tell whoever arrives between 3:00 pm and 5:00 pm to bring their bikes into the garage.

Conclusion

The Google smart home is getting a major upgrade with the addition of Gemini to its Nest cameras and smart home experience. With AI-generated summaries, automations, and a Gemini-powered search feature, users will have a more intelligent and intuitive smart home experience.

FAQs

Q: What is Gemini?
A: Gemini is a set of artificial intelligence (AI) large language models (LLMs) developed by Google.

Q: What features does Gemini bring to the Google smart home?
A: Gemini brings AI-generated summaries, automations, and a Gemini-powered search feature to the Google smart home.

Q: Is Gemini available for all Nest cameras?
A: No, Gemini is only available for cameras in a Nest Aware Plus subscription.

Q: Will Gemini work with Nest Aware’s person identification tech?
A: Not yet, but Google is working on it.

Q: When will the updated Google Assistant be available?
A: The updated Google Assistant is coming soon and will be available for existing Nest smart speakers and displays.

Three from MIT named 2024-25 Goldwater Scholars | MIT News

0

MIT students Ben Lou, Srinath Mahankali, and Kenta Suzuki have been selected to receive Barry Goldwater Scholarships for the 2024-25 academic year. They are among just 438 recipients from across the country selected based on academic merit from an estimated pool of more than 5,000 college sophomores and juniors, approximately 1,350 of whom were nominated by their academic institution to compete for the scholarship.

Since 1989, the Barry Goldwater Scholarship and Excellence in Education Foundation has awarded nearly 11,000 Goldwater scholarships to support undergraduates who intend to pursue research careers in the natural sciences, mathematics, and engineering and have the potential to become leaders in their respective fields. Past scholars have gone on to win an impressive array of prestigious postgraduate fellowships. Almost all, including the three MIT recipients, intend to obtain doctorates in their area of research.

Ben Lou

Ben Lou is a third-year student originally from San Diego, California, majoring in physics and math with a minor in philosophy.

“My research interests are scattered across different disciplines,” says Lou. “I want to draw from a wide range of topics in math and physics, finding novel connections between them, to push forward the frontier of knowledge.”

Since January 2022, he has worked with Nergis Mavalvala, dean of the School of Science, and Hudson Loughlin, a graduate student in the LIGO group, which studies the detection of gravitational waves. Lou is working with them to advance the field of quantum measurement and better understand quantum gravity.

“Ben has enormous intellectual horsepower and works with remarkable independence,” writes Mavalvala in her recommendation letter. “I have no doubt he has an outstanding career in physics ahead of him.”

Lou, for his part, is grateful to Mavalvala and Loughlin, as well as all of his scientific mentors that have supported him along his research path. That includes MIT professors Alan Guth and Barton Zwiebach, who introduced him to quantum physics, as well as his first-year advisor, Richard Price; current advisor, Janet Conrad; Elijah Bodish and Roman Bezrukavnikov in the Department of Mathematics; and David W. Brown of the San Diego Math Circle.

In terms of his future career goals, Lou wants to be a professor of theoretical physics and study, as he says, the “fundamental aspects of reality” while also inspiring students to love math and physics.

In addition to his research, Lou is currently the vice president of the Assistive Technology Club at MIT and actively engaged in raising money for Spinal Muscular Atrophy research. In the future, he’d like to continue his philanthropy work and use his personal experience to advise an assistive technology company.

Srinath Mahankali

Srinath Mahankali is a third-year student from New York City majoring in computer science.

Since June 2022, Mahankali has been an undergraduate researcher in the MIT Computer Science and Artificial Intelligence Laboratory. Working with Pulkit Agrawal, assistant professor of electrical engineering and computer science and head of the Improbable AI Lab, Mahankali’s research is on training robots. Currently, his focus is on training quadruped robots to move in an energy-efficient manner and training agents to interact in environments with minimal feedback. But in the future, he’d like to develop robots that can complete athletic tasks like gymnastics.

“The experience of discussing research with Srinath is similar to discussions with the best PhD students in my group,” writes Agrawal in his recommendation letter. “He is fearless, willing to take risks, persistent, creative, and gets things done.”

Before coming to MIT, Mahankali was a 2021 Regeneron STS scholar, which is one of the oldest and most prestigious awards for math and science students. In 2020, he was also a participant in the MIT PRIMES program, studying objective functions in optimization problems with Yunan Yang, an assistant professor of math at Cornell University.

“I’m deeply grateful to all my research advisors for their invaluable mentorship and guidance,” says Mahankali, extending his thanks to PhD students Zhang-Wei Hong and Gabe Margolis, as well as assistant professor of math at Brandeis, Promit Ghosal, and all of the organizers of the PRIMES program. “I’m also very grateful to all the members of the Improbable AI Lab for their support, encouragement, and willingness to help and discuss any questions I have,”

In the future, Mahankali wants to obtain a PhD and one day lead his own lab in robotics and artificial intelligence.

Kenta Suzuki

Kenta Suzuki is a third-year student majoring in mathematics from Bloomfield Hills, Michigan, and Tokyo, Japan.

Currently, Suzuki works with professor of mathematics Roman Bezrukavnikov on research at the intersection of number and representation theory, using geometric methods to represent p-adic groups. Suzuki has also previously worked with math professors Wei Zhang and Zhiwei Yun, crediting the latter with inspiring him to pursue research in representation theory.

In his recommendation letter, Yun writes, “Kenta is the best undergraduate student that I have worked with in terms of the combination of raw talent, mathematical maturity, and research abilities.”

Before coming to MIT, Suzuki was a Yau Science Award USA finalist in 2020, receiving a gold in math, and he received honorable mention from the Davidson Institute Fellows program in 2021. He also participated in the MIT PRIMES program in 2020. Suzuki credits his PRIMES mentor, Michael Zieve at the University of Michigan, with giving him his first taste of mathematical research. In addition, he extended his thanks to all of his math mentors, including the organizers of MIT Summer Program in Undergraduate Research.

After MIT, Suzuki intends to obtain a PhD in pure math, continuing his research in representation theory and number theory and, one day, teaching at a research-oriented institution.

The Barry Goldwater Scholarship and Excellence in Education Program was established by U.S. Congress in 1986 to honor Senator Barry Goldwater, a soldier and national leader who served the country for 56 years. Awardees receive scholarships of up to $7,500 a year to cover costs related to tuition, room and board, fees, and books.

Unreal Engine 5.5: Top 5 Features

0
  1. Mobile Game Development

Unreal Engine, in most people’s estimations, has always played second fiddle to Unity in the realm of mobile game development. Epic Games is determined to address this in their latest version 5.5 of Unreal Engine.

Much of the new additions centre around the Mobile Forward Renderer that boasts greater visual fidelity on the platform. New features include D-buffer decals, rectangular area lights, capsule shadows, moveable IES textures for points and spotlights, volumetric fog, and Niagara particle lights. These broad improvements help strengthen Unreal Engine’s offering.

Improvements have also been made to the Mobile Previewer, which is ideal for mobile games content development. Practically speaking, this means users can both capture and preview specific Android device profiles.

  1. Rendering Improvements

No Unreal Engine update would be complete without a slew of rendering engine improvements. Lumen continues to go from strength to strength with the new ability to run at 60Hz on platforms for which there is hardware support. All of this is made possible thanks to improved hardware ray tracing (HWRT).

Additionally, Unreal Engine 5.5 has a better-than-ever path tracer, which is finally production-ready. That means we’ve got a wider range of support features, including new additions like sky atmosphere and volumetric clouds. It’s also now got Linux support.

Alongside rendering but always related are changes to Substrate, Unreal Engine’s material authoring framework. What began experimentally in version 5.2 is now in beta. Artists will need to continue caution when using it in a production context but it’s not far off being ready, especially when working on nonlinear material production.

  1. Animation Sequencer

Sequencer, Unreal Engine’s nonlinear animation editor, receives some much-needed attention. Alongside better filtering, the interface is more intuitive with settings that are far easier to control.

Animation layers have been made non-destructive which is a significant step forward in content management that so many animators demand. Layers can be set to additive or override as well as weighted against each other. This makes it much easier to fine-tune animations without having to start all over.

And finally, we love the new ability to trigger cinematic scenarios based on player choices during gameplay. This level of control has never been seen before and helps to strengthen Unreal Engine’s animation support.

  1. Animation Control Rig

Very much related to Sequencer, animation improvements continue with the animation control rig. In main, that means animators can now apply contact deformation and better cartoon-style squash-and-stretch for more realistic animation effects. Animation deformers can now easily be applied to characters in Sequencer with no trouble at all.

For animators who love an easy life, Unreal Engine now ships with an Animator Kit plugin which contains a set of ready-made Control Rigs with built-in deformers. Alongside this, the Modular Control Rig moves to Beta with a number of UI and UX improvements.

  1. VR Scouting

Virtual Scouting was introduced in the previous release, Unreal Engine 5.4, so it hasn’t been around for long. It’s still very much a work in progress but it’s showing a lot of promise. This time around, we have new opportunities for customisation via an extensive API, a new VR Content Browser, and a Transform Gizmo that is customisable via Blueprint. All of this strengthens the virtual Scouting toolset as a ready-to-go out-of-the-box experience.

For more information head over to the Unreal Engine 5 blog, where Epic details every new feature.

Conclusion

Unreal Engine 5.5 is a significant update that brings numerous improvements to the table. With its focus on mobile game development, rendering, animation, and VR scouting, it’s clear that Epic Games is committed to making Unreal Engine the go-to engine for game developers.

FAQs

Q: What are the new features in Unreal Engine 5.5?
A: The new features include improvements to mobile game development, rendering, animation, and VR scouting.

Q: What are the improvements to mobile game development?
A: The improvements include a new Mobile Forward Renderer, D-buffer decals, rectangular area lights, capsule shadows, moveable IES textures for points and spotlights, volumetric fog, and Niagara particle lights.

Q: What are the improvements to rendering?
A: The improvements include a better-than-ever path tracer, Lumen running at 60Hz on platforms with hardware support, and Linux support.

Q: What are the improvements to animation?
A: The improvements include a more intuitive interface for Sequencer, non-destructive animation layers, and the ability to trigger cinematic scenarios based on player choices during gameplay.

Q: What are the improvements to VR scouting?
A: The improvements include new opportunities for customisation via an extensive API, a new VR Content Browser, and a Transform Gizmo that is customisable via Blueprint.

Old School Telcos Revive with AI

Unlock the Editor’s Digest for free

Roula Khalaf, Editor of the FT, selects her favourite stories in this weekly newsletter.

A Legacy Telco’s Revival

From the distressed debt dustbin to an AI darling, all in less than a year. Not long ago, US regional telecoms company Lumen Technologies was seen as a corpse that had been scavenged by Wall Street’s most vicious vulture funds. Now it is the toast of the AI party. Its shares have gained more than 700 per cent since the spring. With its market capitalisation approaching $10bn, suddenly its near $20bn debt load does not look so imposing.

New Tech Needs Old Pipes

New tech increasingly needs old pipes, and a bunch of previously moribund companies is trumpeting a revival. At Lumen, the rally is being driven by $8bn worth of contracts that it has signed with the likes Google, Amazon and Meta for so-called private connectivity fabric. Lumen, formed last decade through the combination of Level 3 and CenturyLink, is an internet backbone provider. AI hyperscalers and datacentres are buying the access to connect to its ultrafast fibre network.

Frontier Communications’ Potential

Another supposedly stagnant legacy telco may also be able to benefit from the renewed importance of old-school connectivity. It may help Frontier Communications to secure a sweeter buyout price from Verizon than the current $20bn on offer. The former’s shareholders believe that dull infrastructure suddenly has more importance in a world where data loads are accelerating and Big Tech has the deep pockets to pay toll collectors to ensure that their demands are met.

Lumen’s Debt Restructuring

The cash from deals with Big Tech is advanced up front while the accounting revenue recognition is more conservative. Interest expense and capital spending remain enormous at Lumen while much of the business is mature and declining — think copper wires.

A Complex Debt Restructuring

A complex debt restructuring completed earlier this year gave Lumen increased liquidity as well as a reprieve from pending debt maturities. Still, with annual free cash flow of less than $2bn and a business otherwise in what seemed like terminal decline, an eventual bankruptcy filing had looked inevitable.

Conclusion

Lumen’s management has leaned hard into the AI hype, declaring that their pipes are essential to unleashing the power of an emerging game-changing technology. Whether that turns out to be true or not, Lumen and peers are well positioned to raise capital while the sun shines.

Frequently Asked Questions

Q: What is Lumen Technologies?
A: Lumen Technologies is a US regional telecoms company that provides internet backbone services.

Q: Why is Lumen’s stock price increasing?
A: Lumen’s stock price is increasing due to its recent contracts with Big Tech companies, including Google, Amazon, and Meta, for private connectivity fabric.

Q: What is the significance of Lumen’s debt restructuring?
A: Lumen’s debt restructuring has given the company increased liquidity and a reprieve from pending debt maturities, allowing it to focus on its future growth.

Q: Is Lumen’s management optimistic about its future?
A: Yes, Lumen’s management is optimistic about its future, citing the importance of its pipes in the emerging AI technology landscape.

Dropbox Lays Off 20% of Staff

0

Layoffs at Dropbox

Dropbox is laying off 528 employees in a move that will reduce its global workforce by 20 percent, CEO Drew Houston announced today.

Background

Houston wrote that Dropbox’s core file sync and sharing “business has matured, and we’ve been working to build our next phase of growth with products like Dash,” an “AI-powered universal search” product targeted to business customers. The company’s “current structure and investment levels” are “no longer sustainable,” according to Houston.

Reasons for Layoffs

“We continue to see softening demand and macro headwinds in our core business,” Houston wrote. “But external factors are only part of the story. We’ve heard from many of you that our organizational structure has become overly complex, with excess layers of management slowing us down.”

Previous Layoffs

Dropbox previously cut 500 employees in an April 2023 round of layoffs. At the time, Houston said that Dropbox’s business was profitable but growth was slowing.

New Layoffs

Today, Houston said that Dropbox is “still not delivering at the level our customers deserve or performing in line with industry peers. So we’re making more significant cuts in areas where we’re over-invested or underperforming while designing a flatter, more efficient team structure overall.”

Severance Package

In a Securities and Exchange Commission filing, Dropbox said it expects to “make total cash expenditures of approximately $63 million to $68 million in connection with the reduction in force, primarily consisting of severance payments, employee benefits and related costs.” Laid-off employees are eligible for 16 weeks of pay, plus one additional week of pay for each year of tenure, Houston wrote. He also said the laid-off workers “will receive their Q4 equity vest” and will be eligible for a pro-rated payment equivalent to their 2024 bonus target.

Conclusion

In conclusion, Dropbox’s layoffs are a significant move that will impact the company’s global workforce. The decision was made to reduce complexity and inefficiencies, while also focusing on building new products and services to drive growth.

A better way to control shape-shifting soft robots | MIT News

0

Imagine a slime-like robot that can seamlessly change its shape to squeeze through narrow spaces, which could be deployed inside the human body to remove an unwanted item.

While such a robot does not yet exist outside a laboratory, researchers are working to develop reconfigurable soft robots for applications in health care, wearable devices, and industrial systems.

But how can one control a squishy robot that doesn’t have joints, limbs, or fingers that can be manipulated, and instead can drastically alter its entire shape at will? MIT researchers are working to answer that question.

They developed a control algorithm that can autonomously learn how to move, stretch, and shape a reconfigurable robot to complete a specific task, even when that task requires the robot to change its morphology multiple times. The team also built a simulator to test control algorithms for deformable soft robots on a series of challenging, shape-changing tasks.

Their method completed each of the eight tasks they evaluated while outperforming other algorithms. The technique worked especially well on multifaceted tasks. For instance, in one test, the robot had to reduce its height while growing two tiny legs to squeeze through a narrow pipe, and then un-grow those legs and extend its torso to open the pipe’s lid.

While reconfigurable soft robots are still in their infancy, such a technique could someday enable general-purpose robots that can adapt their shapes to accomplish diverse tasks.

“When people think about soft robots, they tend to think about robots that are elastic, but return to their original shape. Our robot is like slime and can actually change its morphology. It is very striking that our method worked so well because we are dealing with something very new,” says Boyuan Chen, an electrical engineering and computer science (EECS) graduate student and co-author of a paper on this approach.

Chen’s co-authors include lead author Suning Huang, an undergraduate student at Tsinghua University in China who completed this work while a visiting student at MIT; Huazhe Xu, an assistant professor at Tsinghua University; and senior author Vincent Sitzmann, an assistant professor of EECS at MIT who leads the Scene Representation Group in the Computer Science and Artificial Intelligence Laboratory. The research will be presented at the International Conference on Learning Representations.

Controlling dynamic motion

Scientists often teach robots to complete tasks using a machine-learning approach known as reinforcement learning, which is a trial-and-error process in which the robot is rewarded for actions that move it closer to a goal.

This can be effective when the robot’s moving parts are consistent and well-defined, like a gripper with three fingers. With a robotic gripper, a reinforcement learning algorithm might move one finger slightly, learning by trial and error whether that motion earns it a reward. Then it would move on to the next finger, and so on.

But shape-shifting robots, which are controlled by magnetic fields, can dynamically squish, bend, or elongate their entire bodies.


The researchers built a simulator to test control algorithms for deformable soft robots on a series of challenging, shape-changing tasks. Here, a reconfigurable robot learns to elongate and curve its soft body to weave around obstacles and reach a target.

Image: Courtesy of the researchers

“Such a robot could have thousands of small pieces of muscle to control, so it is very hard to learn in a traditional way,” says Chen.

To solve this problem, he and his collaborators had to think about it differently. Rather than moving each tiny muscle individually, their reinforcement learning algorithm begins by learning to control groups of adjacent muscles that work together.

Then, after the algorithm has explored the space of possible actions by focusing on groups of muscles, it drills down into finer detail to optimize the policy, or action plan, it has learned. In this way, the control algorithm follows a coarse-to-fine methodology.

“Coarse-to-fine means that when you take a random action, that random action is likely to make a difference. The change in the outcome is likely very significant because you coarsely control several muscles at the same time,” Sitzmann says.

To enable this, the researchers treat a robot’s action space, or how it can move in a certain area, like an image.

Their machine-learning model uses images of the robot’s environment to generate a 2D action space, which includes the robot and the area around it. They simulate robot motion using what is known as the material-point-method, where the action space is covered by points, like image pixels, and overlayed with a grid.

The same way nearby pixels in an image are related (like the pixels that form a tree in a photo), they built their algorithm to understand that nearby action points have stronger correlations. Points around the robot’s “shoulder” will move similarly when it changes shape, while points on the robot’s “leg” will also move similarly, but in a different way than those on the “shoulder.”

In addition, the researchers use the same machine-learning model to look at the environment and predict the actions the robot should take, which makes it more efficient.

Building a simulator

After developing this approach, the researchers needed a way to test it, so they created a simulation environment called DittoGym.

DittoGym features eight tasks that evaluate a reconfigurable robot’s ability to dynamically change shape. In one, the robot must elongate and curve its body so it can weave around obstacles to reach a target point. In another, it must change its shape to mimic letters of the alphabet.

Animation of orange blob shifting into shapes such as a star, and the letters “M,” “I,” and “T.”
In this simulation, the reconfigurable soft robot, trained using the researchers’ control algorithm, must change its shape to mimic objects, like stars, and the letters M-I-T.

Image: Courtesy of the researchers

“Our task selection in DittoGym follows both generic reinforcement learning benchmark design principles and the specific needs of reconfigurable robots. Each task is designed to represent certain properties that we deem important, such as the capability to navigate through long-horizon explorations, the ability to analyze the environment, and interact with external objects,” Huang says. “We believe they together can give users a comprehensive understanding of the flexibility of reconfigurable robots and the effectiveness of our reinforcement learning scheme.”

Their algorithm outperformed baseline methods and was the only technique suitable for completing multistage tasks that required several shape changes.

“We have a stronger correlation between action points that are closer to each other, and I think that is key to making this work so well,” says Chen.

While it may be many years before shape-shifting robots are deployed in the real world, Chen and his collaborators hope their work inspires other scientists not only to study reconfigurable soft robots but also to think about leveraging 2D action spaces for other complex control problems.

AI Agents Arrive

A New Era in AI: Anthropic’s Early Access to Basic AI Agents

A Giant Leap Forward

Last month, Anthropic made a groundbreaking announcement by providing early access to basic AI agents for the masses. This significant development marks a major milestone in the evolution of artificial intelligence (AI), moving beyond the limitations of chatbots and opening up new possibilities for widespread adoption.

Breaking Free from Chatbots

For years, chatbots have dominated the early days of generative AI. While they have been successful in providing basic customer service and answering simple queries, they have been limited in their capabilities and often fell short in providing meaningful interactions. Anthropic’s latest innovation takes a significant leap forward by offering basic AI agents that can engage in more sophisticated conversations and tasks.

What are Basic AI Agents?

Basic AI agents are designed to be more intelligent and capable than traditional chatbots. They are trained on vast amounts of data and can learn from their interactions with humans, allowing them to improve their performance over time. These agents can be integrated into various applications, including customer service, language translation, and content creation.

Key Features and Capabilities

Anthropic’s basic AI agents come equipped with several key features and capabilities, including:

Conversational Intelligence

Basic AI agents can engage in more natural and human-like conversations, understanding context and nuances of language.

Task-Oriented

These agents can perform specific tasks, such as data entry, research, and content creation, freeing up human workers to focus on higher-value tasks.

Continuous Learning

Basic AI agents can learn from their interactions with humans, improving their performance and accuracy over time.

Implications and Opportunities

The availability of basic AI agents has far-reaching implications and opportunities, including:

Increased Productivity

Basic AI agents can automate routine and repetitive tasks, freeing up human workers to focus on more creative and strategic work.

Improved Customer Experience

These agents can provide personalized and efficient customer service, enhancing the overall customer experience.

New Business Models

Basic AI agents can enable new business models and revenue streams, such as AI-powered content creation and language translation services.

Conclusion

Anthropic’s early access to basic AI agents marks a significant milestone in the development of artificial intelligence. This innovation has the potential to transform industries and revolutionize the way we live and work. As AI continues to evolve, it will be exciting to see how these agents are applied and the impact they will have on society.

FAQs

Q: What is the difference between basic AI agents and chatbots?
A: Basic AI agents are more intelligent and capable than traditional chatbots, with the ability to engage in more sophisticated conversations and perform specific tasks.

Q: How do basic AI agents learn and improve?
A: Basic AI agents learn from their interactions with humans, using machine learning algorithms to improve their performance and accuracy over time.

Q: Can basic AI agents be integrated into various applications?
A: Yes, basic AI agents can be integrated into various applications, including customer service, language translation, and content creation.

Q: What are the implications of basic AI agents on the workforce?
A: Basic AI agents have the potential to automate routine and repetitive tasks, freeing up human workers to focus on more creative and strategic work.

It Needs to Feel Like Magic

0

What do Rings of Power, The White Lotus, and Time Bandits Have in Common?

As well as being some of the biggest TV shows of the 2020s, they feature stunningly inventive title sequences – all of which were created by small and tight-knit Seattle-based design company Plains of Yonder.

Different Approach to Title Sequences

Plains of Yonder takes a different approach to larger LA-based firms, fostering close relationships with show-runners in order to create “high aesthetic designs” that turn key creative themes and Easter eggs from the shows into intricate works of art that have the power, as we saw with the wildly popular opening to The White Lotus, to become cultural phenomenons in their own right.

Visualizing Sound: The Concept Behind the Rings of Power Title Sequence

The concept of “visualizing sound” is key to the Rings of Power title sequence. This comes back to season one, when the studio approached Plains of Yonder to pitch and come up with ideas. The series is really going back in time, showing how Tolkien’s world is created by time and the ebb and flow of nations and alliances. They came across his creation myth, which is about these Ainur, these angelic beings singing the world into creation. This connection between music and physics brought them to the idea of Cymatics, which is a real-world phenomenon where you put sand on a plate, and you apply a vibration to that plate, and according to hertz and resonance, you get a certain symmetrical pattern that forms.

Revisiting the Idea for Season 2

For Season 2, there were some challenges. Of course, you need to evolve from Season 1. You always want to push the envelope and go further into a concept, but you need to create a new tone with the exact same music. So, they stayed in the world of Cymatics, with the same music, but created a new tone. Experimentation is the way they work. They experiment and kind of run wild, trying to capture a mood. For every shot that’s in the main title, there’s probably 20 or 30 shots that they built roughly to see if it captures the mood.

Title Sequences: A Puzzle to Solve

Title sequences are often abstract and expressive, but also have to convey practical information and match up with music. Are these constraints challenging or do they present creative opportunities? They see title sequences as a little portal, a way to distill and encapsulate a world, and bring viewers into that world. They like to stay a little bit tangential to the show, so they don’t see their job as replicating the show. They try to stay creative, experiment, and push the envelope.

The Titles for The White Lotus Became a Viral Hit

I think you’re surprised when every dance club in the world is playing the song. That you can’t predict, but as a social satire, it hit something for people. They do think that when you have a show and a main title and music, and everything is in real synergy, it can be culturally powerful.

What’s the Secret to Stopping Viewers from Hitting ‘Skip Intro’?

If they’ve done their job right, it feels integral to the show. When they think about titles that are super important and critical to them, there’s certain shows you have where you would never skip it, because it gets you excited in a way that if you didn’t have it, you would feel incomplete. That’s electric.

Conclusion

Plains of Yonder’s approach to title sequences is unique, fostering close relationships with show-runners and creating “high aesthetic designs” that turn key creative themes and Easter eggs into intricate works of art. Their concept of “visualizing sound” and experimentation with Cymatics have resulted in stunning title sequences that have captured the attention of audiences worldwide.

FAQs

Q: What is the concept of “visualizing sound”?

A: The concept of “visualizing sound” is the idea of using real-world phenomena, such as Cymatics, to create a visual representation of music and sound.

Q: How do you approach revisiting an idea for a second season?

A: They experiment and try to capture a mood, building multiple shots to see if it captures the mood they’re going for.

Q: What’s the secret to stopping viewers from hitting ‘Skip Intro’?

A: If they’ve done their job right, it feels integral to the show, and gets viewers excited and engaged.

Q: How do you stay creative and push the envelope in title sequences?

A: They experiment, try new things, and stay true to their creative vision.

Q: What’s the importance of synergy between the show, title sequence, and music?

A: When everything is in real synergy, it can be culturally powerful and capture the attention of audiences worldwide.