メタバースは失敗しました。少なくとも、2021年に語られた「次のインターネット」は、まだ私たちの日常になっていません。
The metaverse failed. At least, the next internet promised in 2021 has yet to become part of everyday life.
当時の構想は壮大でした。画面の向こう側を見るのではなく、その中へ入り、遠くにいる誰かと同じ場所にいる感覚を共有する。Metaはそれを「身体性を持つインターネット」と呼びました。しかし現実に広がったのは、重いヘッドセットと、会議室やイベント会場を模した、どこか寂しい仮想空間でした。
The vision was grand. Instead of looking through a screen, we would step inside it and share the feeling of occupying the same place as someone far away. Meta called it an embodied internet. What reached the public instead were heavy headsets and strangely lonely virtual spaces modeled on conference rooms and event halls.
Founder’s Letter, 2021(Meta)
構想そのものが間違っていたのでしょうか。
Was the vision itself wrong?
Omochi編集部は、そう考えていません。足りなかったのは、世界へ入るための入口だけではなく、世界を作り、動かし続けるための技術だったのではないでしょうか。建物も、道路も、物体も、そこで起きる出来事も、人間が一つずつ設計する。利用者が何かをすれば、開発者があらかじめ書いた反応が返ってくる。その方法では、現実のように広く、予測不能で、絶えず変化する世界を作れません。
Omochi’s editors do not think so. What was missing may have been more than an entrance into the world. It was the technology to create that world and keep it alive. Every building, road, object, and event had to be designed by hand. When a user did something, the world returned a response a developer had scripted in advance. That approach cannot produce a world as broad, unpredictable, and continuously changing as reality.
いま、Google DeepMind、Meta、NVIDIA、World Labs、Runwayが競う「世界モデル」は、この限界を崩し始めています。仮想世界を完成品として用意するのではなく、人やAIの動きを受け取り、次に起きる世界をその場で生成する。これは動画生成AIの次の機能ではありません。メタバースの生産方式そのものを変える技術です。
The world models now being pursued by Google DeepMind, Meta, NVIDIA, World Labs, and Runway are beginning to break that constraint. Instead of supplying a virtual world as a finished product, they take the actions of people and AI as input and generate what happens next on the spot. This is not simply the next feature in generative video. It changes the production system of the metaverse itself.
※ 世界モデルとは、環境の内部表現を学習し、行動や出来事の結果として次に何が起きるかを予測するAIのこと。映像として描き出す場合もあれば、抽象的な状態の変化として予測する場合もある。
※ A world model is an AI that learns an internal representation of an environment and predicts what happens next as a result of actions and events. Some render the prediction as video; others predict abstract state changes.
「作られた世界」から「生成され続ける世界」へ
From a Built World to a World That Keeps Generating Itself
これまでのゲームやメタバースでは、世界の大部分は制作物でした。壁は壁として配置され、ドアには開閉の動作が設定され、登場人物には会話の選択肢が与えられます。自由に見えても、体験できることの多くは誰かが先回りして作ったものです。
In conventional games and metaverse platforms, most of the world was an authored asset. Walls were placed as walls, doors were assigned opening animations, and characters were given dialogue choices. However free the experience appeared, much of what a user could do had been built in advance by someone else.
世界モデルは、この順序を逆にします。
World models reverse that order.
人が右へ進めば、右側に見える景色を生成する。ロボットがコップへ手を伸ばせば、腕とコップの次の状態を予測する。雨を降らせれば、路面や視界がどう変わるかを描く。世界を保存された舞台装置ではなく、行動に応じて次の状態を計算する仕組みへ変えていきます。
When a person moves right, the model generates the view to the right. When a robot reaches for a cup, it predicts the next state of the arm and the cup. When rain begins, it renders the resulting changes to the road and visibility. The world becomes less a stored stage set and more a system that calculates its next state in response to action.
Google DeepMindのGenie 3は、その分かりやすい例です。テキストから作った環境を720p、毎秒24フレームで動かし、利用者の操作に応じて次の映像を生成します。一度作った動画を再生するのではなく、進む方向が変わるたびに、その先の世界を作り直しているのです。
Google DeepMind’s Genie 3 offers a clear example. It runs environments generated from text at 720p and 24 frames per second, producing the next image in response to user input. It is not replaying a video made earlier. Each change in direction prompts it to create the world ahead anew.
Genie 3: A new frontier for world models(Google DeepMind)
ここに、かつてのメタバースとの決定的な違いがあります。古いメタバースでは、作り手が世界を完成させてから利用者を招きました。世界モデルでは、利用者が動くこと自体が、世界の続きを作る入力になります。
That is the decisive difference from the earlier metaverse. Its creators finished a world and then invited users inside. With a world model, the act of moving through the environment becomes an input that creates what comes next.
現実の動きを戻せば、ループは閉じる
Feed Real Movement Back, and the Loop Closes
では、現実の人間やロボットの動きをリアルタイムで世界モデルへ戻せば、仮想世界と現実をつなげられるのでしょうか。
Could a virtual world then be connected to reality simply by feeding the movements of a person or robot back into the model in real time?
原理的には、その通りです。
In principle, yes.
カメラやセンサーが現在の状態を観測する。世界モデルが「この動きをしたら何が起きるか」を予測する。人やロボットが実際に動く。その結果を再び観測し、予測と現実のずれを修正する。観測、予測、行動、再観測が途切れずに回れば、モデルの中の世界は現実からフィードバックを受けながら更新され続けます。
Cameras and sensors observe the current state. A world model predicts what will happen if a particular movement is made. The person or robot acts. The result is observed again, allowing the model to correct the gap between prediction and reality. If observation, prediction, action, and renewed observation continue without interruption, the modeled world can keep updating itself through feedback from reality.
これはまだ思考実験ではありません。MetaのV-JEPA 2を使った実験では、ロボットが現在の状態と目標を見て、複数の行動候補がもたらす結果をモデルの中で予測します。そして最も目標へ近づく動きを一つ実行し、新しい状態を観測した次の瞬間には、もう一度計画を立て直します。
This is no longer only a thought experiment. In a Meta experiment using V-JEPA 2, a robot observes its present state and its goal, then predicts the consequences of several candidate actions inside the model. It executes the movement most likely to approach the goal and, as soon as it observes the new state, plans again.
※ モデル予測制御とは、モデルを使って複数の行動の結果を先に予測し、最適な行動を少しだけ実行した後、最新の観測を使って再び計画する制御手法のこと。
※ Model-predictive control uses a model to forecast the results of several actions, executes a small portion of the best one, then plans again using the latest observation.
Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning(Meta)
人間に置き換えれば、身体の動きや視線を読み取り、仮想世界がそれに反応し、その反応を見た人間が次に動くということです。かつて「アバターを仮想空間へ表示する」ことが中心だったメタバースは、ここで「人間と生成世界が互いを更新する」ものへ変わります。
For a human, the equivalent would be a virtual world reading bodily movement and gaze, responding to them, and eliciting the person’s next movement in turn. A metaverse once centered on placing an avatar in a virtual space becomes something in which a person and a generated world continuously update one another.
各社が作っているのは、同じ世界の別々の器官
Each Company Is Building a Different Organ of the Same World
ただし、現在の世界モデルが一つでこの循環を完成させているわけではありません。「世界モデル」という名前の下で、各社は別々の難題を解いています。
No world model completes this cycle by itself today. Under the common label of world models, companies are solving different parts of the problem.
Genie 3やRunwayのGWM-1が得意とするのは、行動に反応する映像です。入力を受けながら次のフレームを作り続けることで、人がその場にいるように感じる視覚的な入口を作ります。
Genie 3 and Runway’s GWM-1 specialize in imagery that responds to action. By generating each successive frame from live input, they create the visual entrance that makes a person feel present.
Introducing Runway GWM-1(Runway)
World LabsのMarbleは、振り返っても同じ場所が残る、探索可能な3D空間を重視します。世界が毎秒美しく描き直されても、さっき置いた物が消えたり、部屋の形が変わったりすれば、そこを「場所」とは感じられません。メタバースに必要なのは生成能力だけでなく、過去を引き継ぐ能力です。
World Labs’ Marble emphasizes explorable 3D spaces that remain in place when the user turns around. However beautifully a world redraws itself, it will not feel like a place if an object set down moments ago disappears or the shape of a room changes. The metaverse needs more than generative capacity. It needs the ability to preserve its past.
Marble: A Multimodal World Model(World Labs)
MetaのV-JEPA 2が担うのは、見栄えではなく行動の選択です。未来を一枚ずつ描く代わりに、どの行動が目標へ近づくかを抽象的な表現の中で比べます。NVIDIAのCosmosは、こうしたモデルをロボットや自動運転の開発に使うため、合成データ、モデルのカスタマイズ、シミュレーションをまとめた産業基盤を狙います。
Meta’s V-JEPA 2 handles the choice of action rather than visual presentation. Instead of painting the future frame by frame, it compares in an abstract representation which action will move closest to a goal. NVIDIA’s Cosmos seeks to become the industrial layer for deploying such models in robotics and autonomous driving, combining synthetic data, model customization, and simulation.
※ 合成データとは、実世界で収集する代わりに、シミュレーションや生成モデルを使って人工的に作られた学習データのこと。現実では収集が危険・高価・困難な状況を、大量に用意できる。
※ Synthetic data is training data produced artificially with simulation or generative models instead of being collected in the real world. It can supply, at scale, situations that are dangerous, expensive, or hard to capture in reality.
NVIDIA Makes Cosmos World Foundation Models Openly Available to Physical AI Developer Community(NVIDIA Blog)
World Labsは、この状況を「レンダラー」「シミュレーター」「プランナー」という三つの機能に整理しています。人に見せる映像を作る機能、物理的に一貫した状態を計算する機能、次の行動を選ぶ機能です。いま見えている競争は、どの一社の映像が最も美しいかではありません。この三つの器官を、誰が一つの循環へ接続できるかという競争です。
World Labs divides the field into three functions: renderers, simulators, and planners. One produces imagery for people to see, another calculates a physically coherent state, and the third selects the next action. The contest is not over which company can produce the prettiest picture. It is over who can connect these three organs into a single cycle.
A Functional Taxonomy of World Models(World Labs)
| 会社 | モデル | 器官 | 役割 |
|---|---|---|---|
| Google DeepMind | Genie 3 | レンダラー | 操作に反応する世界を映像として生成 |
| Runway | GWM-1 | レンダラー | 環境・アバター・ロボットのリアルタイム映像 |
| World Labs | Marble | シミュレーター | 過去が残る、探索可能な3D空間 |
| Meta | V-JEPA 2 | プランナー | 行動の結果を予測し、次の一手を選ぶ |
| NVIDIA | Cosmos | 産業基盤 | 合成データと開発ツールで物理AIを支える |
| Company | Model | Organ | Role |
|---|---|---|---|
| Google DeepMind | Genie 3 | Renderer | Generates a world as video that responds to input |
| Runway | GWM-1 | Renderer | Real-time imagery for environments, avatars, robots |
| World Labs | Marble | Simulator | Explorable 3D spaces that keep their past |
| Meta | V-JEPA 2 | Planner | Predicts outcomes of actions and picks the next move |
| NVIDIA | Cosmos | Industrial base | Synthetic data and tooling for physical AI |
美しい夢が、再び壊れる場所
Where the Beautiful Dream Could Break Again
ここまで聞けば、メタバースの復活はすぐそこに見えます。現実の動きを読み取り、生成世界へ返す。それだけでループは閉じるように思えます。
At this point, the metaverse can seem close to revival. Read movement from reality and feed it into a generated world. That appears to close the loop.
しかし、ループが閉じることと、その中に信頼できる世界が生まれることは別です。
But closing the loop is not the same as creating a trustworthy world inside it.
反応が少し遅れるだけで、人はそこにいる感覚を失います。コップが机をすり抜ければ、ロボットは間違った物理を学びます。同じ部屋へ戻ったときに配置が変わっていれば、空間は記憶を持てません。二人が同じ物を同時に動かしたとき、それぞれに別の結果が見えれば、共有世界は成立しません。
A small delay in response can destroy a person’s sense of presence. If a cup passes through a desk, a robot learns the wrong physics. If the layout changes when someone returns to the same room, the space has no memory. If two people moving the same object see different results, there is no shared world.
Genie 3も、維持できる一貫性は数分間で、操作の種類、複数の独立したエージェントの相互作用、実在する場所の正確な再現には制約があると説明しています。World Labsも、見た目を作るレンダリング、構造と物理を保つシミュレーション、行動を決める計画を一つのモデルに統合することを、未解決の中心課題に挙げています。
Genie 3 itself says its consistency lasts for minutes and acknowledges constraints on the range of actions, interactions among independent agents, and accurate reproduction of real locations. World Labs likewise identifies the unification of visual rendering, structurally and physically coherent simulation, and action planning as a central open problem.
映像として本物らしいことと、世界として正しいことは違います。生成AIは前者を急速に進歩させました。メタバースの第二幕を決めるのは、後者です。
Looking real on screen and being correct as a world are different things. Generative AI has advanced the former at extraordinary speed. The latter will determine the metaverse’s second act.
最初の住人は、人間ではないかもしれない
The First Residents May Not Be Human
だからこそ、この新しいメタバースは、以前とは違う場所から始まる可能性があります。
That is why this new metaverse may begin somewhere very different from the last one.
最初のブームは、人間を仮想空間へ連れていこうとしました。新しいブームでは、まずロボットとAIエージェントがそこへ入ります。倉庫で荷物を落とす。道路で判断を誤る。家庭で見たことのない道具に触れる。現実では危険で、遅く、高価な失敗を、生成された世界なら何百万回でも繰り返せるからです。
The first boom tried to bring people into virtual space. In the new one, robots and AI agents may enter first. They can drop packages in warehouses, make bad decisions on roads, and encounter unfamiliar tools in homes. Failures that are dangerous, slow, and expensive in reality can be repeated millions of times in a generated world.
※ AIエージェントとは、目標を与えられると、環境を観測しながら自ら行動を選んで実行するAIのこと。チャットで応答するだけでなく、一連のタスクを自律的に進める。
※ An AI agent is an AI that, given a goal, observes its environment and chooses and executes actions on its own. Rather than only replying in chat, it carries a sequence of tasks forward autonomously.
彼らがそこで身につけるのは、仮想世界だけで通用するゲームの攻略法ではありません。モデルの予測を現実のセンサー情報で修正しながら、現実で動くための能力です。世界モデルが先に巨大な産業を作るとすれば、交流イベントや仮想店舗より、ロボット訓練、自動運転、デジタルツインかもしれません。
What they learn there will not be a game strategy that works only in a virtual world. They will learn how to act in reality while correcting model predictions with real sensor data. If world models create a major industry first, it may emerge from robot training, autonomous driving, and digital twins rather than social events or virtual storefronts.
※ デジタルツインとは、工場や都市といった現実の対象を、センサーデータと同期させながらコンピュータ上に再現した仮想の複製のこと。現実を止めずに、実験や将来の予測ができる。
※ A digital twin is a virtual replica of a real-world subject, such as a factory or a city, kept in sync with sensor data. It allows experiments and forecasts without stopping the real thing.
その先で、技術は再び人間へ戻ってきます。眼鏡が見ている場所を理解し、視線や手の動きに応じて情報を置く。遠くにいる人の動きを受け取り、自分と同じ空間にいるかのように反映する。現実の部屋と生成された世界が、一方通行の映像ではなく、互いに変化を返し合うようになる。
Beyond that point, the technology returns to people. Glasses understand the place in view and position information in response to gaze and hand movement. They receive the movements of someone far away and reproduce them as if that person occupied the same space. A real room and a generated world cease to be a one-way image and begin returning changes to each other.
Omochi編集部の見立てでは、世界モデルはメタバースの復活です。ただし、復活するのは「VRの中に店を作れば人が集まる」という2021年の熱狂ではありません。
Omochi’s view is that world models represent a revival of the metaverse. What returns, however, will not be the 2021 conviction that opening a store in VR would automatically bring a crowd.
人間が手作業で作った世界へ入るのが、第一幕でした。人間とAIが動くたびに世界が続きを生成し、現実からのフィードバックで自らを修正する。それが第二幕です。
In the first act, people entered worlds built by hand. In the second, the world generates what comes next whenever people or AI move, then corrects itself through feedback from reality.
メタバースは死んだのではありません。世界を作り終えてから公開するという発想が、先に限界を迎えたのです。次の世界は、完成を待っていません。私たちが一歩を踏み出した瞬間、その足元から作られ始めます。
The metaverse did not die. The idea that a world had to be finished before it could be opened reached its limit first. The next world will not wait for completion. The moment we take a step, it will begin forming beneath our feet.
こうした次の産業を巡る話をじっくり交わす場所探しは、Omochiにお任せください。
Let Omochi find the right setting for a deeper conversation about the industries taking shape next.