Skip to content

前Meta科学家希望将视觉人工智能引入工厂车间| TechCrunch

2 9 月, 2026

人工智能正在改变我们周围的一切,但到目前为止,它在很大程度上仍然被数字领域所遏制。然而,越来越多的初创公司希望将其带入现实世界。

由两位前Meta研究科学家创办的初创公司就是这样一家公司。该公司成立于2024年11月,致力于开发前沿视觉模型,旨在帮助机器更有效地与物理环境进行交互。

本周,该公司推出了其最新型号Isaac 0.5 ,其创造者称其旨在为机器提供在工业环境中“感知、推理和行动”的能力。具体而言,该软件能够帮助视觉引导机器人在仓库或工厂车间等复杂环境中导航。它还可以帮助公司从这些机器人录制的视频中提取视觉智能。

Isaac 0.5也作为开放式权重模型发布,因此任何人都可以检查其参数和培训材料。

这家初创公司由Armen Aghajanyan和Akshat Shrivastava共同创立,他们之前曾在Meta的人工智能研究部门Fundamental AI Research ( FAIR )工作。两人将他们的软件视为工业自动化部署的未来。

该公司表示: “今天的物理AI迫使人们做出错误的选择:通才基础模型需要为每个实例提供多个专用云GPU ,或者处理感知或控制的狭窄模型,但从未两者兼而有之。”

Aghajanyan and Shrivastava say their tool is unlike existing models in the space because it is general-purpose, meaning that it’s not built for one specific, repetitive task. Instead, they say, the model is designed to be flexible depending on the particular environment (or situation) it is in.

In an interview, Shrivastava asked me to consider what goes into a simple physical process like organizing boxes: “Imagine there’s a robot being deployed to sort packages right now. What are the tasks it would need to do?”

Such a relatively simple task indeed consists of many steps. A robot would first have to read the label on the package, do some spatial analysis to understand where the boxes are, and decide which one to pick up. If it’s picking up a series of boxes, it would have to plan which boxes to pick up and in which order.

Perceptron’s software is designed to help robots find their way through each step of the process. To be clear, the industry already has software that can help machines do most of those tasks, but there are few programs that are designed to do it flexibly.

Where does the data for this algorithmic alchemy come from?

Models like Isaac 0.5 learn operational skills by ingesting gargantuan amounts of video training data. Perceptron says its new model was fed on a million hours of what is known as general video to teach its algorithm to identify particular settings, visuals, and scenarios. The company also relied heavily on what is known as

— video captured, typically through a GoPro or a wearable camera, from the perspective of a person completing a physical task — as well as

, which are similarly used to teach AI systems movements by recording repetitive human actions.

While Perceptron isn’t disclosing the sources of its training data, Shrivastava said that the company had “internally built petabyte-scale datasets that span across modalities, whether it’s images, text, video, etc. all the way through robotic trajectories.”

The utility of a software that can help robots operate competently in warehouses is obviously vast, and Perceptron thinks it’s well-positioned to lead that wave of automation. The startup is ready to market its software to a variety of vendors, and thus potentially see its intelligence layer integrated into a broad array of industries.

Those industries include manufacturing, logistics and warehousing, security, mobility, as well as media and entertainment.

“Nothing like this really exists out there,” said Aghajanyan. “We’re really excited about it.”

The company previously raised $21 million from Bessemer Venture Partners, Foundation Capital, and S32, according to Pitchbook. SmartGateVC also participated in the company’s founding round. TechCrunch understands the startup is in the process of closing an additional round.

This story has been updated to include additional funding details.