Connect with us

Tech

OpenAI’s Hugging Face breach has reignited the debate over alignment and control

Last week, an unreleased model built by OpenAI breached Hugging Face’s systems during internal testing, and a lot of theoretical research suddenly became very practical. The hack was the first verifiable case of an AI lab losing control of its own model, chaining together exploits to gain access it never should have had. But while the AI industry has been united in its alarm, a split has emerged in how researchers want to respond.

For some, the problem is a basic cybersecurity issue: The sandbox failed to contain the model, and Hugging Face’s cybersecurity systems failed to keep it out. Those problems can be solved by patching bugs and building more robust control and containment methods for increasingly capable AI that is prone to go rogue in autonomous environments. 

But another camp takes a more pessimistic view. For them, AI’s rapidly increasing capabilities mean that trying to control rogue models is a losing game. The only robust security comes from making sure the models aren’t trying to escape in the first place — a challenge often referred to as alignment. In alignment terms, the problem is that OpenAI’s model was trying to cheat, and solving that problem is more urgent than short-term containment efforts.

Judging by its public statements, OpenAI is taking both camps seriously. The company has rushed to patch the bugs involved in the hack, and it referenced both alignment and monitoring approaches in its statement after the breach became public. But the company’s response also suggests a philosophy that has left many safety researchers alarmed: Rather than slowing down or stopping the development of more capable models, it should instead focus on building stronger cages around them.  

“As models take on longer and more complex tasks, failures that evaluations miss may carry greater consequences,” OpenAI said in a postmortem of the incident. “We will keep working to narrow the gap between evaluation and deployment: testing models over longer trajectories, improving alignment, building monitoring that can intervene, and giving users clearer visibility and control.”

OpenAI’s latest frontier model is more likely than its predecessor to engage in misaligned behaviors. Image Credits:OpenAI

There’s also reason to think OpenAI’s models are becoming less aligned as they become more powerful. According to OpenAI’s system card ,GPT-5.6 Sol is significantly more prone to agentic misalignment than its predecessor, GPT-5.5. In deployment simulations, the company also found the model was more likely to circumvent restrictions, engage in destructive actions, and perform unauthorized data transfers than GPT-5.5. Those figures were largely overlooked on first release, but in the wake of the breach, they’re getting a second look — particularly since Sol was one of the models involved.

In a social media post, OpenAI’s Head of Strategic Futures, Dean Ball, argued that monitoring and transparency were the best ways to keep those tendencies in check. 

“These issues will become more salient as the capabilities of models improve, and as the stakes of their deployment grow,” he said. “The solution is neither alarmism nor complacency. Instead, I believe the solution lies in careful measurement and monitoring, an engineering mentality, and transparency.”

One former OpenAI researcher told TechCrunch that the firm tends to focus on “outer alignment” rather than “inner alignment” — essentially the difference between an AI system that understands a set of values and can represent them convincingly, and one that actually has those values at its core. In this case, outer alignment wasn’t enough to convince the model that it shouldn’t cheat on the test.

OpenAI did not respond to repeated requests for more information.

For alignment-focused researchers, OpenAI’s response isn’t good enough. Zvi Mowshowitz, a writer who focuses on new AI developments, argued that OpenAI’s decision to treat the incident as an infrastructure problem may help solve the immediate cybersecurity issues, but it will fail in the long term. 

“This is an alignment problem,” Mowshowitz wrote in a recent Substack blog. “This is the models being misaligned, and all of the OpenAI models showing severe signs of exactly the problem we are all most worried about, in a way that is likely embedded into their training on a deep level. The entire training pipeline needs to be addressed in this light, or it will only get worse.”

Several experts told TechCrunch that the incident is evidence that today’s training methods produce systems that optimize for outcomes rather than internalize human intentions. 

Redwood Research, a nonprofit AI safety and security research organization, classified OpenAI’s model behavior in this case as “score-seeking misalignment,” a pattern in which AI models try to get a high score regardless of instructions, side effects, or downstream consequences. 

“Models with these alignment properties could set up a ‘Potemkin village’ of false successes to make it look like things are fine when they’re not,” Alex Mallen and Girish Gupta, two researchers at Redwood, wrote in a recent paper

Score-seeking behavior and other misalignment isn’t unique to OpenAI. Anthropic has published several papers on emergent misalignment behaviors that surface when its frontier models are optimized or placed in autonomous environments, including deception, reward-hacking, and malicious autonomy

“We still consistently see models trying to circumvent constraints and act deceptively when they are asked to do tasks at the edge of their abilities,” Neev Parikh, an AI safety researcher at alignment nonprofit METR, told TechCrunch via email. “In our frontier risk report, we saw this behavior fairly consistently, despite efforts from companies to try and reduce this behavior.”

Implicit in OpenAI’s response to the Hugging Face incident is the assumption that development will continue on even more capable systems, whether they are suitably aligned at their core or not. Going back to the drawing board isn’t really an option when the business models of AI firms depend on delivering the next generation of models. If it may never be possible to know with certainty that a model is fully aligned, then the practical question comes down to how to safely contain and control increasingly capable systems.

“There’s not yet a good understanding of how to align the most capable AI systems, but there’s much more consensus about how to control them,” Steven Adler — a former safety researcher at OpenAI and current chief scientist of Guidelight AI Standards, an organization that publishes a standard for avoiding incidents like the Hugging Face one — told TechCrunch. “Every company has a ways to go in achieving this.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

source

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Tech

Are brain waves the next unlock for physical AI?

The frontier of physical AI is a Jenga game in a warehouse in San Leandro, California.

That warehouse is occupied by Encord, a company that builds data tooling used to train AI models. Andrew Ceja is a pilot — the company’s term for its robotic trainers — and he’s carefully pulling wooden blocks from a tottering tower while wearing a headset with a camera that tracks what he sees. That alone is fairly common for collecting robot training data, but this headset includes sensors that measure his brain waves as he carefully disassembles the block tower.

Encord is one of a growing number of startups betting the next real constraint on humanoid and warehouse will be the scarcity of real-world physical training data, and which is building a business not just to manage that data but to manufacture it.

The brain wave headset Ceja is wearing was built by Zander Labs, a German neuroscience startup that’s betting measuring brain activity — to deduce mental states like error, intent, and surprise — can create a more useful dataset to train models. Encord’s work with Zander is currently a trial run; Encord says the goal is to build an initial brain wave-tagged dataset, run it through customer robotics models, and evaluate if it actually improves performance before deciding whether to scale it up.

Lukas Gehrke, a Zander neuroscientist supervising the work, says that the amount of brain activity used at any point during a given task offers clues for model builders trying to figure out when they need to deploy their highest-effort models.

This is the “bleeding edge” of the effort to solve the robotics data bottleneck, according to Vineeth Velmurugan, Encord’s head of robot learning. A veteran of OpenAI’s robot lab and Berkshire Grey, the warehouse automation firm, Velmurugan joined Encord to build the company’s internal data-creation team.

Encord was founded to help companies building machine-vision applications annotate data and evaluate models. As their customers — Velmurugan says they work with many leading robotics firms but that he’s not authorized to name them — began to apply end-to-end learning to robotic manipulation tasks, executives realized they would have to produce training data themselves, rather than simply manage it. “The data simply does not exist,” Velmurugan said.

The bet that generative AI can do for robots what it’s done for chatbots keeps running into this same wall. Self-driving car companies collect physical-world data themselves, but that’s hard to scale. Training from video can work, but it lacks the fidelity of real-world data. Velmurugan says it will take a dataset something like five times the size of YouTube’s video corpus to break through — a scale that helps explain why data-generation itself has become a business and not just a research problem.

Feed your egocentric data needs

Companies building robot brains are now turning to two main sources: “egocentric” video collected by workers wearing cameras, often augmented with additional camera angles and other metrics, and data from robots operated remotely. Encord does both, drawing egocentric data from several factories around the globe, and using its San Leandro facility to experiment with new modalities, like brain waves, or collect datasets around specific skills for fine-tuning.

When TechCrunch visited, pilots were using leader-follower rigs — paired robotic arms, one controlled directly by a human operator and one that mimics its movements — to create data about tasks like pouring coffee from a pot into mugs (very sloshy) and stacking poker chips. “Every humanoid company has asked us for these pieces,” Velmurugan says.

Storage racks held cartons of fake flowers in vases, books, plastic vegetables, kitty litter trays and scoops, bags and bundles of wires — the stock in trade for training manipulators for household tasks.

At one of these stations, another pilot, Sofia Infante, maneuvers robotic arms to plug and unplug ethernet cables from the back of a server — the kind of work data center operators would love to be automated, if only robots could manipulate them with the required precision. Taking a spin behind the controls, I was able to see why that’s still out of reach: Pincers are far less dexterous than human fingers and lack the degrees of freedom we take for granted in our arms.

Another new data modality that Encord is developing uses a set of sensors strapped to the forearm to detect electrical signals in muscles. Video taken of human hands manipulating objects typically doesn’t capture the entire hand, but Velmurugan hopes to build a 3D depiction of where the hand is at any time based on the arm sensors, creating a more robust understanding for models.

Encord’s datasets are annotated with physical descriptions of what each video contains — “right hand tightens bolt” — to aid LLM-based models in understanding what is happening. Velmurugan estimates this kind of dense annotation is worth 100 times as much as “junky ego data” for training specific tasks, and it only costs 20 times more to produce, which is a good trade, on paper.

But “20 times more” is still real money, and that’s the catch: Scraping text off the internet, the way LLM makers built their models by pulling from Stack Overflow and the rest of the web, cost frontier labs next to nothing. Generating physical training data does not, and that’s the limit of the physical-AI-as-LLM comparison. This kind of data has to be manufactured, not just collected, and that changes the economics of building these models.

Velmurugan says that progress is being made — with Encord’s visibility into programs across the industry, he’s able to see startups and frontier labs alike figure out what works and what doesn’t to improve physical AI models. That vantage point — sitting between many robotics companies at once — is also part of Encord’s pitch. It can spot which data techniques are gaining traction industrywide before any single customer can.

That will keep the dozen or so pilots at Encord’s facility busy. Both Infante and Ceja are part of a burgeoning workforce developing the building blocks for neural networks; they previously worked at Scale, another AI data annotation firm, before joining Encord.

Ceja had worked at a waste management company where his interest in technology found him in charge of keeping a robotic trash sorter in good working order. Now, as the Jenga tower topples, he says he enjoys the challenge of solving training tasks for robots — “It’s something new every day!”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

source

Continue Reading

Tech

Europe got its own TBPN-style live show, and everyone’s angling for a guest spot

The European answer to TBPN is here and ready to go live five days a week, starting July 27. 

Luke Knight and Ronan Chambers launched the London-based European Technology Network (ETN) last October, breaking down tech trends and news during a two-day-a-week livestream. Right now, the show is livestreamed on X and YouTube and has garnered more than 5 million views.

On Monday, the network announced a $1.6 million seed round from top players in the media ecosystem, including Powerhouse Capital, Axel Springer SE (which owns Business Insider and Politico), one of the co-founders of the popular media publication LADbible, and angel investors from OpenAI and DeepMind. With this fresh capital, the network is announcing its largest expansion yet. 

It’s now moving into a big studio in Kings Cross (where all the hot London AI startups are situated), expanding the team (right now of just eight), launching a newsletter, and, starting today, moving into a five-day-a-week live-show schedule, which will soon see Knight and Chambers interview the likes of George Robson (a partner at Sequoia) and Rishi Sunak (former U.K. prime minister and senior advisor to Anthropic and Microsoft). 

Speaking to TechCrunch, Knight and Chambers said ETN has already become a hot stop on the press tour for European startups — they’ve spoken to the founder of Synthesia, the CFO of Legora, the founder of Granola, and Kanishka Narayan, the U.K.’s first AI minister. They’ve even had American investors stop by the show when they are in town, including one from Andreessen Horowitz.

“ETN was born out of a gaping hole in the industry,” Chambers told TechCrunch. “It’s centered around pace.” He said the current media ecosystem in the U.K. cannot keep up with how fast the tech scene is moving. For example, so far this year, London startups have raised $14.7 billion according to Dealroom. Six companies have raised more than $500 million: Wayve, Superintelligence, ElevenLabs, Recursive, Ineffable Intelligence, and Isomorphic Labs, the latter three of which were founded by DeepMind alumni.

“These are things that have never happened in Europe before,” Chambers continued, referring to the speed at which capital is flowing through the ecosystem. As the show became more popular, Chambers said they were getting around 70 pitches a week from guests looking to come on the show. They would try to cram 12 interviews into two hours, twice a week, but eventually it got to be too much. “We needed an outlet that could move at the pace of that,” he said of both the interest in the show and how fast Europe’s tech ecosystem is moving, “which is the reason we’re going from two days a week to five days a week.” 

The five-day format will look quite similar to the two-day format. There will be a live show from 12 p.m. U.K. time to 3 p.m., breaking down trending stories, and then for two hours, they will have guests on the show talking about whatever they want. Chambers said they also want to start hosting debates, roundtables, and a Shark Tank-style pitching session on the show.

“We want to make it as useful as we can for the ecosystem,” Chambers said. “There needs to be more discourse around AI in Europe. There needs to be more discourse around venture capital and cash flowing into the ecosystem. There needs to be more discourse around the amazing things that are happening in the tech ecosystem, and our role is to be the stage in which people can shout about all the amazing things they’re doing.” 

The show makes its money from ad dollars, like most media publications, and big-name sponsors already include prediction market Polymarket, blockchain company Base, and the AI audio darling ElevenLabs. 

When asked about the influence TBPN has had on them, Knight and Chambers said they indeed do look at John Coogan and Jordi Hays, founders of TBPN (which recently sold to OpenAI for what some say was a nine-figure sum), as pioneers of this new tech media ecosystem. The show became a place for tech guests to appear and chat with friendly faces, announcing new product releases, hires, or funding news. “My thinking was, if we can have an ITV and a BBC, why wouldn’t we have a regional version of this?” Chambers continued. 

Europe is a big place, though, with more than 40 countries and over 200 languages spoken (24 of which are recognized by the European Union). Chambers said that although ETN will report from London, he and Knight are making an effort to bring on guests from across the continent. Aside from bringing guests into the studio, they also travel to the hottest tech conferences around Europe. For example, they’ve broadcast from the Panathēnea Conference in Athens and from inside the Louvre in Paris for the RAISE AI Summit. 

“You have all these different cultures, these different minds coming together and creating different products,” Chambers said. “You get a taste of what makes [Europe] a superpower.” 

With all this, they said they would never turn an American founder away should they want to come on the show. “It’s a European technology network, but we think there’s massive [global] opportunities,” Chambers said. Knight added to that, noting how often conversations pit the European tech ecosystem against that of the U.S. 

“We are globally optimistic,” Knight said. “We are pushing global prosperity from Europe. Wherever you want to go and build your company, wherever is the best place to go and build that company, go and do that, and we will shout for you to go and do that.” 

He and Chambers also don’t necessarily see themselves as journalists; rather, they consider themselves tech insiders curious about what is going on and why. They also don’t see themselves as replacing traditional media and instead intend to work in tandem with those publications. “We rely on traditional media,” Chambers said, adding that is how they find much of the news that they talk about on ETN. 

Overall, the duo hopes to help document the stories coming from the new wave of European success, from ElevenLabs in London to Lovable in Stockholm, to help the upcoming generation understand that technology is one way to drive a nation forward.

Discussing the impact of European success stories, Chambers said, “It riles up the next generation to the point where it’s no longer cool to finish university and go into banking or consulting. People want to leave university and go straight into building a startup, and I think that’s an amazing thing.”

This piece was updated to clarify who invested in the company and when the show was coming out.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

source

Continue Reading

Tech

Snapchat now lets you share what you’re listening to in real time

Snapchat is now letting users share the music they’re listening to in real time via a new feature in Snap Map, TechCrunch has exclusively learned.

Called “Now Playing,” the feature lets users link their streaming service account (starting with Spotify) and share what they’re currently listening to via Snap Map, and see the songs their friends are enjoying.

The feature has been integrated with Spotlight videos, which are Snapchat’s TikTok-like short-form videos, too: You can save any songs you discover in a Spotlight video directly to Spotify, or visit the track’s page on the platform to do so.

Users can choose who can see their listening activity, and the feature won’t show your last-played song or listening history, though sharing will remain active as long as you have opened Snapchat within the past 24 hours. If a user is inactive for more than 24 hours, sharing will be paused automatically, and only resumed when the app has been opened again. You can also pause sharing for three hours, 24 hours, or indefinitely.

The new feature comes as social media platforms continue to embrace music discovery and sharing. TikTok, where viral trends often shape global music charts, lets users share songs from streaming services to the social network, and also save songs they come across. Instagram, meanwhile, allows users to share what they’re listening to through Notes, which are the short status updates that appear at the top of users’ DM inboxes.

Even Spotify itself has leaned into social music sharing, launching a feature that allows users to share what they’re listening to with their friends in real time.

“Music is one of the most personal ways people express themselves, and it becomes even more meaningful when it brings friends closer,” Manny Adler, Snapchat’s head of music, said in an emailed statement. “Now Playing adds a new layer of expression and discovery to Snap Map, helping Snapchatters share the soundtrack to their day and find new music through the people they already know.”

The integration adds another feature to Snap Map, which has more than 450 million monthly users. Launched in 2017, Snap Map was initially meant to be a way for users to see their friends’ locations and browse public Snaps from around the world. Over time, the feature has expanded to include local hotspots, activities, and now, music discovery.

Snapchat says the new feature is rolling out to users in regions where both Spotify and Snapchat are available, with Canada coming soon.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

source

Continue Reading