咖啡空间成具身大脑训练场?WRC世界模型热度高涨:技术路线仍未收敛,能力徘徊于“学步”与“青春期”之间Is coffee space a training ground for embodied brain? The WRC world model is gaining popularity: the technological roadmap has not yet converged, and the ability is hovering between "learning to walk" and "adolescence"
“2026世界机器人大会”(WRC)于8月19日在北京亦庄启幕。在今年的WRC展厅里,有一处特殊的咖啡空间:柜台内有人类咖啡师,柜台上有长程手冲机械臂协助制作咖啡,柜台外则有轮式机器人穿梭在人群中为顾客提供服务。
不同于刚刚在北京银河SOHO落地的京东旗下七鲜咖啡24小时智能化无人咖啡店,WRC展厅里的这个咖啡空间更想做“人机共生”。这是成立仅一年多、今年上半年完成数亿美元融资的具身智能公司——无界动力为其具身大脑挑选的训练场。
为什么是咖啡空间?
无界动力创始人兼CEO(首席执行官)张玉峰在展厅现场接受包括《每日经济新闻》记者(以下简称“每经记者”)在内的媒体采访时回答很务实:咖啡店场景的任务和物品SKU(单品)相对集中,不像家庭环境那样千差万别——每家的床单、衣服、茶杯都不一样。但咖啡店又有足够的背景泛化:光线在变、人流在变、顾客留在桌上的私人物品在变,机器人必须识别并避开它们,收纸巾时不能误收手机,放咖啡时要选一个合适的位置。
在张玉峰看来,这种场景“业务边界明确、环境变量动态”,既能让公司自研的隐空间世界模型在稳定的业务框架内持续学习,又必须应对海量真实变化,从而检验其将已有技能迁移至不同人员、位置与状态的能力。
每经记者注意到,能够理解、预测物理世界的世界模型,如今正是行业方兴未艾的技术焦点,也是本届WRC的热点之一。但当下行业尚未进入“技术收敛”阶段,咖啡空间背后,是各家技术与商业化路线的同台竞技。
“咖啡场景”成机器人落地热门试验田
走进本届WRC展厅,最直观的感受是:机器人不再只是在舞台上表演、在拳击台上对抗,而是纷纷亮出在各种场景里干活的真本领。其中,“做咖啡”俨然成了检验机器人真实落地能力的热门“试验田”。
深圳安诺机器人是饮品机器人细分领域唯一拿下独立展位的企业,同时也亮相合作方京东的展台。每经记者在现场看到,其双臂具身机器人拉花咖啡印花吧和单臂AI(人工智能)拉花咖啡印花亭,通过高精度机械臂、视觉识别和机器学习,可迅速完成一杯带精美拉花的咖啡,用户还可上传AIGC(人工智能生成内容)生图选择独一无二的印花图案。
在擎朗与挪瓦咖啡联合呈现的“机器人咖啡馆”中,人形机器人XMAN-R1以“特聘咖啡师”身份亮相。该机器人基于擎朗自研通用VLA(视觉—语言—动作)架构训练出的岗位化垂域模型KEENON ProS。XMAN-R1能自主完成取杯、咖啡萃取等步骤,并将制作好的咖啡放置拉花机供客人DIY印花。
与京东七鲜咖啡主打的全流程无人化、封闭贩售机式咖啡机器人不同,无界动力的咖啡空间里仍有人类咖啡师,机器人也不待在封闭空间或吧台后,而是成了咖啡空间里的服务员,从真人咖啡师手中接过咖啡端给顾客、收空杯、清理桌面,甚至在扔完垃圾后主动给双手做紫外线消杀。
“商业服务场景中,顾客通常期待看到人,很少有咖啡店品牌想变成完全无人的咖啡店,所以人机共生是核心。”张玉峰强调:“我们不再把机器人围起来,而是让它真正与人配合。”

无界动力展区的“人机共生沉浸式机器人咖啡空间” 图片来源:每经记者 郑欣蔚 摄
每经记者在现场注意到,真人咖啡师头上佩戴着无界动力自研的AnySense EGO第一视角多模态数据采集设备,用于采集真人操作数据,反哺后续模型升级。在张玉峰看来,硬科技的迭代需要与场景深度结合、与B端客户深入结合,找到“沿途下蛋”的机会,用真实场景数据牵引高质量数据反哺,形成数据飞轮,推动具身大脑持续迭代。
每经记者还了解到,在商业服务场景上,无界动力从咖啡店切入,接下来4个月将在北京与一家咖啡店品牌继续合作试运营,并逐渐转入常态化运营,计划今年四季度进入韩国市场;咖啡店之后,快餐店、酒店是下一步目标。工业场景上,无界动力的机器人已与汽车零部件企业、能源企业展开合作,真正进入产线。
“例如折纸盒子,北京的外资车厂对此有诉求,希望机器人帮忙折盒子、整理零部件。我们就打造了泛工业、泛家庭场景下能折纸盒的机器人,做到了‘一机多能’。大部分展示中,机器人可能只做一件事,但我们的叠袜子、叠衣服其实是一个模型,把物品丢给它,它就会挨个处理。”张玉峰说。
在展区现场,无界动力还展示了三台机器人基于同一模型,围绕折叠、分类、包装等环节实现自主分工与协同。泛化能力的缺失,也是宇树科技创始人王兴兴在大会演讲中提及的制约机器人走进生活、走进家庭的最大瓶颈。“泛化”同样是张玉峰反复提及的关键词,“一模多能”的通用化能力是他认为具身大脑需要努力的方向。
世界模型技术路线仍未收敛
在“AI教母”李飞飞创办的World Labs(世界实验室)带动下,热度高涨的世界模型路线在本届WRC上呈现出更大声势:北京人形机器人创新中心发布具身多模态大一统模型PelicanUnify,大晓机器人带来开悟世界模型,极佳视界展示新一代具身基模GigaBrain-0.7⋯⋯押注世界模型路线的参展商纷纷在本届WRC上秀出自己的具身大脑。
在超维动力展台,一台全尺寸人形机器人KAIBot与现场观众展开火热的乒乓球对打。面对速度快、落点难测的来球,机器人在毫秒级窗口内完成识别、判断与击球。支撑这一能力的,是其自研的KAI世界模型。该模型围绕“生成虚拟世界进行动作交互—模型评估理解—掌握物理世界规律”的三步闭环构建,让机器人在进入真实对抗前,已在仿真世界中完成海量预训练,摸透乒乓对局的操作。
在今年WRC主论坛的演讲中,宇树科技创始人王兴兴同样透露,宇树科技早在2020年年初就开始探索基于视频生成的世界模型方向,中间因效果不理想而搁置,直到去年又重启投入,“我们公司对AI模型这块的投入一直非常大,而且应该也是目前公司资金和人力投入最大的方向”。
同样押注世界模型的无界动力选择的是“隐空间世界模型+强化学习”技术路线,张玉峰毫不掩饰对这条路线的信心,甚至用了“最接近终局”的表述。
这位曾深耕汽车行业多年、出任过地平线副总裁的创业者以自动驾驶为例,对比了当下流行的VLA(视觉—语言—动作)路线:“车不压路牙子,是因为它见过百万、千万次不压路牙子的轨迹,而不是像人一样真正理解那是路牙子以及压上去的后果。”张玉峰指出,VLA本质上还是模仿学习,而非对时空关联性、物理世界因果性的真正理解。
对于VLA与世界模型的路线之争,张玉峰坦言:“今天没有一个路线是物理AI或AGI(通用人工智能)的终极路线,自动驾驶也没有到最终AGI路线,否则L4规模为什么没起来?”他表示,希望能看到行业在算法、模型上有更大创新。他同时透露,无界动力将在9月发布更大参数量、不依赖任何已有开源模型的新版本基座模型。
此外,张玉峰还指出,算法、硬件、场景是机器人产业迭代的三位一体,但目前整个产业链的硬件成熟度、标准化仍有待提升。目前人形机器人市场规模仍小,这对硬件成熟度打磨、成本下降、一致性提升都是挑战。“中国供应链确实是最强的,但不代表完全成熟。我们在‘有和无’上非常强,但在成本和一致性上,整个行业还有很多路要走。”
对于目前具身智能行业的发展阶段,张玉峰则判断:“具身智能目前的能力大概处于学步期与青春期之间,能做一些基础工作,但太复杂的事情还不能一下子就学会。”但他很喜欢本届大会开幕视频中反复提及的主题——“再试一次”。
“机器人从以前的跑步到现在在工厂里干活,总会有失败的时候,就像人从小到大成长一样,再试一次就会成功。”在张玉峰的构想中,具身智能最终会像新的水电煤、新算力基础设施一样,从实现当下的价值创造,到采集更多数据燃料与反馈,形成更强的能力,成为新型基础设施。
免责声明:本文内容与数据仅供参考,不构成投资建议,使用前请核实。据此操作,风险自担。
每日经济新闻
The 2026 World Robot Conference (WRC) kicked off on August 19th in Yizhuang, Beijing. In this year's WRC exhibition hall, there is a special coffee space: there are human baristas inside the counter, long-range hand drawn robotic arms on the counter to assist in making coffee, and wheeled robots outside the counter shuttle through the crowd to provide services to customers.
Unlike the 24-hour intelligent unmanned coffee shop under JD's Seven Fresh Coffee, which has just landed in Beijing Galaxy SOHO, the coffee space in the WRC exhibition hall is more focused on "human-machine symbiosis". This is the training ground selected by Wujie Power, a embodied intelligence company that has only been established for over a year and completed hundreds of millions of dollars in financing in the first half of this year, for its embodied brain.
Why is it a coffee space?
Zhang Yufeng, founder and CEO of Wujie Power, answered pragmatically during an interview with media including Daily Economic News reporters at the exhibition hall: the tasks and SKUs of items in the coffee shop scene are relatively concentrated, unlike the vast differences in the home environment - each family's bed sheets, clothes, and tea cups are different. But coffee shops also have enough background generalization: the lighting is changing, the flow of people is changing, and the personal items left by customers on the table are changing. Robots must recognize and avoid them, and when collecting tissues, they cannot accidentally collect mobile phones. When placing coffee, they must choose a suitable location.
In Zhang Yufeng's view, this scenario has clear business boundaries and dynamic environmental variables, which allows the company's self-developed hidden space world model to continuously learn within a stable business framework, while also having to deal with massive real changes, thus testing its ability to transfer existing skills to different personnel, locations, and states.
Every journalist has noticed that world models that can understand and predict the physical world are now the emerging technological focus of the industry and one of the hot topics of this year's WRC. But the current industry has not yet entered the stage of "technological convergence", and behind the coffee space is the competition of various technologies and commercialization routes.
Coffee scene becomes a popular experimental field for robot landing
Entering the exhibition hall of this year's WRC, the most intuitive feeling is that robots are no longer just performing on stage or fighting on the boxing ring, but are showcasing their true abilities to work in various scenes. Among them, "making coffee" has become a popular "test field" for testing the real landing ability of robots.
Shenzhen Anno Robotics is the only company in the beverage robot segment that has won an independent booth, and also appeared at the booth of its partner JD.com. Every time the reporter saw it on site, their dual arm embodied robot latte art coffee printing bar and single arm AI latte art coffee printing booth could quickly complete a cup of coffee with exquisite latte art through high-precision robotic arms, visual recognition, and machine learning. Users could also upload AIGC (Artificial Intelligence Generated Content) images to select unique printing patterns.
In the "Robot Cafe" jointly presented by Qinglang and Novacoffee, the humanoid robot XMAN-R1 appeared as a "specially appointed barista". The robot is based on the KEENON ProS, a specialized vertical domain model trained on the self-developed general VLA (vision language action) architecture by Qinglang. XMAN-R1 can independently complete the steps of cup picking, coffee extraction, etc., and place the prepared coffee on the latte machine for customers to DIY print.
Different from the whole process unmanned, closed vending machine coffee robot featured by JD Seven Fresh Coffee, there are still human baristas in the coffee space of boundless power, and the robot does not stay behind the closed space or the bar, but has become a waiter in the coffee space, taking the coffee end from the real baristas to customers, collecting empty cups, cleaning the table, and even taking the initiative to do ultraviolet disinfection and sterilization for both hands after throwing garbage.
In commercial service scenarios, customers usually expect to see people, and few coffee shop brands want to become completely unmanned coffee shops, so human-machine symbiosis is the core. "Zhang Yufeng emphasized," We no longer surround robots, but let them truly cooperate with people
The "Human Machine Symbiosis Immersive Robot Coffee Space" in the Boundless Power Exhibition Area. Photo source: Photo by journalist Zheng Xinwei
Every time the reporter noticed on site, the real barista was wearing the AnySense EGO first view multimodal data acquisition device developed by Boundless Power on their head, which was used to collect real operation data and provide feedback for subsequent model upgrades. In Zhang Yufeng's view, the iteration of hard technology needs to be deeply integrated with scenarios and B-end customers, finding opportunities to lay eggs along the way, using real scenario data to drive high-quality data feedback, forming a data flywheel, and promoting continuous iteration of embodied brains.
According to the reporter, in terms of commercial service scenarios, Wujie Power has entered the coffee shop market. In the next four months, it will continue to cooperate with a coffee shop brand in Beijing for trial operation and gradually transition to normal operation. It plans to enter the Korean market in the fourth quarter of this year; After the coffee shop, fast food restaurants and hotels are the next targets. In the industrial scene, boundaryless robots have cooperated with automotive parts companies and energy companies to truly enter the production line.
For example, for origami boxes, foreign-funded car factories in Beijing have demands for robots to help fold boxes and organize parts. We have created robots that can fold paper boxes in industrial and household scenarios, achieving 'one machine, multiple functions'. In most exhibitions, robots may only do one thing, but our stacking socks and clothes is actually a model. If we throw the items to it, it will handle them one by one, "said Zhang Yufeng.
At the exhibition site, Boundless Dynamics also showed that three robots based on the same model realized independent division and collaboration around folding, classification, packaging and other links. The lack of generalization ability is also the biggest bottleneck that restricts robots from entering daily life and households, as mentioned by Wang Xingxing, the founder of Yushu Technology, in his speech at the conference. "Generalization" is also a key word repeatedly mentioned by Zhang Yufeng. The generalization ability of "the first mock examination with multiple abilities" is the direction that he thinks the embodied brain needs to work hard.
The world model technology roadmap has not yet converged
Driven by the World Labs founded by AI godmother Li Feifei, the highly popular world model route presented even greater momentum at this year's WRC: the Beijing Humanoid Robot Innovation Center released the embodied multimodal unified model PelicanUnify, Daxiao Robot brought the enlightened world model, and the excellent vision showcased the new generation of embodied basic model GigaBrain-0.7... Exhibitors who bet on the world model route showcased their embodied brains at this year's WRC.
At the Super Dimensional Power booth, a full-size humanoid robot KAIBot engaged in a heated table tennis match with the live audience. Faced with a fast and unpredictable ball, the robot completes recognition, judgment, and hitting within a millisecond window. Supporting this capability is its self-developed KAI world model. The model is built around a three-step closed-loop of "generating a virtual world for action interaction - model evaluation and understanding - mastering the laws of the physical world", allowing the robot to complete massive pre training in the simulation world and understand the operation of ping-pong matches before entering real confrontation.
In his speech at this year's WRC main forum, Wang Xingxing, founder of Yushu Technology, also revealed that Yushu Technology had been exploring the direction of world models based on video generation as early as the beginning of 2020. However, due to unsatisfactory results, it was put on hold until it resumed investment last year. "Our company has always invested heavily in AI models, and it should be the direction with the largest capital and manpower investment at present.
The unbounded power that also bet on the world model chooses the "hidden space world model+reinforcement learning" technology route, and Zhang Yufeng does not hide his confidence in this route, even using the phrase "closest to the end".
This entrepreneur, who has been deeply involved in the automotive industry for many years and has served as the Vice President of Horizon Robotics, used autonomous driving as an example to compare the popular VLA (Visual Language Action) route: "The car does not press the curb because it has seen millions or millions of trajectories without pressing the curb, rather than truly understanding that it is a curb and the consequences of pressing it like a human." Zhang Yufeng pointed out that VLA is essentially imitation learning, rather than a true understanding of spatiotemporal correlations and physical world causality.
Regarding the debate between VLA and world models, Zhang Yufeng candidly stated: "Today, there is no ultimate path for physical AI or AGI (General Artificial Intelligence), and autonomous driving has not reached the ultimate AGI path. Otherwise, why hasn't L4 scale emerged? ”He expressed his hope to see greater innovation in algorithms and models within the industry. He also revealed that Wujie Power will release a new version of the base model in September with larger parameters and without relying on any existing open source models.
In addition, Zhang Yufeng also pointed out that algorithms, hardware, and scenarios are the three in one iteration of the robotics industry, but the hardware maturity and standardization of the entire industry chain still need to be improved. At present, the market size of humanoid robots is still small, which poses challenges to hardware maturity polishing, cost reduction, and consistency improvement. The Chinese supply chain is indeed the strongest, but it does not mean it is completely mature. We are very strong in 'having and not having', but in terms of cost and consistency, the entire industry still has a long way to go
Regarding the current development stage of the embodied intelligence industry, Zhang Yufeng judged that "the current ability of embodied intelligence is probably between the toddler stage and adolescence, able to do some basic work, but too complex things cannot be learned at once." However, he likes the theme repeatedly mentioned in the opening video of this conference - "Try again.
Robots have gone from running in the past to working in factories now, and there will always be times of failure, just like how humans grow from childhood to adulthood. Trying again will lead to success. "In Zhang Yufeng's vision, embodied intelligence will eventually become a new type of infrastructure, like new water, electricity, coal, and new computing infrastructure, from realizing current value creation to collecting more data, fuel, and feedback, forming stronger capabilities and becoming a new type of infrastructure.
Disclaimer: The content and data in this article are for reference only and do not constitute investment advice. Please verify before use. Based on this operation, the risk is borne by oneself.
Daily Economic News