Sep 29, 2026 · Virtastic
Past the 4 GB Wall: Moving a Native 3D Engine to 64-bit WebAssembly
A field report on compiling OpenMW and its whole dependency stack to wasm64: the bugs that printed nothing, the measurements with their conditions, and what it costs.
Want to see it first? Play Morrowind in your browser. The single-player modes run entirely on your machine. Source: github.com/Virtastic/openmw-web, GPL-3.0-or-later.
Abstract
WebAssembly’s 32-bit memory model caps a browser tab at 4 GB. Large native software, such as CAD, simulation and open-world games, hits that ceiling. The Memory64 feature removes it, and Chrome, Edge and Firefox now ship it. Very little has been written about what moving a real codebase onto it involves.
We compiled OpenMW, an open-source engine for the game Morrowind, with its full dependency stack to wasm64. It runs in a desktop Chromium browser with no install, with server-authoritative multiplayer. This paper reports what broke, why each failure was hard to see, what the fixes were, and what the result costs.
Four findings matter beyond this project:
- The failures are silent. Most of our worst bugs printed no error. A black screen, a frozen page whose header still read 60 fps, and a game with no audio for months were all ordinary outcomes.
- Memory64 is a free audit for pointer bugs. Code that survived on 32-bit luck throws a type error under wasm64. That is unpleasant on day one and valuable on day two.
- Threading is the real port. With no Asyncify, every blocking call had to become inline work, a polled state, or a spin on shared memory. We hit five separate deadlocks.
- The price is measurable. Our wasm64 build used about 27% more CPU per frame than wasm32 in our test, and the browser build ran at about 1.2 times native CPU per frame in a controlled comparison. Both figures come with conditions, which we state.
The numbers at a glance
- 5.65 MB: engine binary, brotli (33.7 MB raw)
- 1.20×: native CPU time per frame in the browser (5.87 ms vs 4.89 ms, one scene, see §5)
- +27%: CPU per frame, wasm64 vs wasm32 (0.785 vs 0.620 ms median, 3 runs each)
- 16 GiB: heap ceiling, the V8 maximum; a retail client uses about 1.5 GB
- 48 fps: 64 avatars in one area, multiplayer client, tiered detail
- ≈ 2×: fewer streaming stall milliseconds after tuning (2,684 → 1,380 ms)
Measured on the developer’s Windows laptop (RTX 4080 Laptop GPU, Chrome on ANGLE/D3D11), retail Morrowind data, unless stated. Section 5 lists conditions and caveats for every figure.
1. Why this matters to a company
Autodesk took AutoCAD, a codebase about thirty years old, to the web through Emscripten and WebAssembly. Figma reported load times more than three times faster after switching to it. Both chose to compile the C++ they had instead of rewriting it. The business case is not new. What is new is that a tab can now hold a working set larger than 4 GB, which is what those products’ larger customers need.
The alternative for a heavy 3D application has been streaming rendered frames from a GPU server. Vendors and blog estimates put that anywhere from tens of cents to about a dollar per user-hour, before bandwidth. Client-side WebAssembly moves that cost onto the user’s own device. The trade is what this paper measures: extra CPU, a large first download, and a narrower set of browsers.
Three groups can reuse what we built:
- Simulation and visualisation teams on OpenSceneGraph (flight and marine simulation, GIS, medical visualisation, osgEarth).
- Robotics and machine-learning teams on Bullet (PyBullet, gym-pybullet-drones, panda-gym), who want a shareable physics scene without a server.
- Engineering-software vendors with a large native codebase and a customer who cannot install anything.
2. What we built
One statically linked openmw.wasm, built with Emscripten 6.0.1, with threads, SIMD, WebAssembly exceptions and a 1.5 GiB initial heap that can grow to 16 GiB. The full dependency stack was rebuilt from a clean checkout on 2026-08-24.
play/streamfs.js.| Library | What it does here | What it took |
|---|---|---|
| OpenSceneGraph 3.6.5 | Scene graph and renderer | A 746-line patch across 15 files (37 hunks): render-to-texture draw buffers, sized renderbuffer formats, packed depth-stencil, MSAA resolve, forced WebGL2 capabilities. X11 is stubbed with exact signatures. |
| MyGUI 3.4.3 | User interface | A 41-line force-included shim restores std::char_traits for 16- and 32-bit characters, which modern libc++ dropped. |
| Bullet 3.25 | Physics | Static, double precision, wasm64. Bullet is unpatched. OpenMW’s build probe for precision is skipped under cross-compilation. |
| FFmpeg 6.1.2 | Bink video, mp3, Vorbis, WAV | Minimal LGPL configuration. Everything is disabled except a short decoder list. No GPL, version-3 or non-free flags. |
| Boost, Recast/Detour, Lua 5.4, LZ4, ICU, others | Utilities, navigation, scripting, compression, text | Built for wasm64 with no library patches. ICU needed its data file supplied by hand (see story 1). |
3. Field notes
Each story below has the same shape: what we saw, why it was hard to find, what it was, and what a porter should take from it. Evidence is in the commit history and in server/docs/STATUS.md.
Story 1. A bare “null function” on every boot
Symptom. Retail boot ran normally through audio, data loading and physics, then died with RuntimeError: null function. No message, no library name.
Why hard. It looked like graphics, so four graphics-shaped hypotheses were tested and each failed. Undefined symbols were not the cause: a relink with -sERROR_ON_UNDEFINED_SYMBOLS=1 found none. A GL entry point missing on the software renderer was not the cause either: a run under ANGLE-over-SwiftShader produced a byte-identical crash at the identical line. Profiling function names came out plausible and wrong, because the flag shifts the function index space.
Cause. The Emscripten ICU port links ICU’s “data supplied elsewhere” stub, and nothing supplied the data. A number-format lookup returned a null pointer and MessageFormat::format called a virtual through it. In WebAssembly a call through null reads a garbage table index and surfaces as a bare null function. The engine reaches it on every boot, through the settings window’s slider labels.
Fix. A twelve-line program against the same three archives reproduced it. The fix stages the 28.5 MB icudt68l.dat into the preload package and calls u_setDataDirectory("/icu"). The link script now fails if the file is absent.
A later attempt to trim that file from 28.6 MB to 1.9 MB was reverted. The trimmed package kept ICU’s locale index for about 800 locales whose data had been removed, so ICU opened one, got nothing, and called a virtual on null. The commit message records that a control build of the previous branch should have been the first test.
Lesson. A bare
null functionis usually a virtual call on a null from a failed resource lookup. Reproduce the suspected library in isolation before blaming the graphics stack.
Story 2. Memory64 turns latent pointer bugs into loud ones
Symptom. The first wasm64 build threw this on every boot:
Cannot convert 32379232 to a BigInt
Why hard. The 32-bit build never showed it, and the message names no function or library.
Cause. Emscripten’s OpenAL does not implement alcGetProcAddress but still advertises the HRTF and pause extensions. With undefined symbols tolerated, those entry points become stubs. On 32-bit, the results were only used behind counts that came back zero, so it survived on luck. Under memory64 the stub call’s Number pointer meets an i64 parameter and throws.
Fix. The engine’s proc-address lookups return null on the web and HRTF is reported absent. That also fixed a latent 32-bit bug: pausing audio when the tab lost focus called an unresolvable stub.
The same class of bug appeared in the JavaScript-to-C boundary. A hand-written exported function that takes a pointer, called with a JavaScript Number, throws a BigInt error on wasm64. Ours was wrapped in a try/catch, so pasting text would have stopped working with nothing logged. Emscripten’s own malloc and free exports keep working because they carry generated glue, which makes a probe built on them misleading. We fixed it with ccall, and wrote a small toy program that reproduces every pointer-crossing form the engine uses, so the failure can be seen before it is far from its cause.
A third case never crashed at all. A change-detection hash truncated an object pointer to 32 bits. With more than 4 GB of heap, two images in different regions alias, and a healthy video would be force-ended after ten seconds. It cannot reproduce on a small heap.
Lesson. Grep for pointer casts to
intanduint32_tbefore you enable a large heap. Treat every BigInt error as a free audit, and give every hand-written JavaScript-to-C pointer crossing toccall.
Story 3. Two pointer models in one build tree
Symptom. Hours into a build: wasm32 object file can't be linked in wasm64 mode. FFmpeg’s configure reported only “C compiler test failed”.
Why hard. In-tree builds (FFmpeg, LZ4, Boost) leave 32-bit objects that make considers up to date. Our first stamp-based guard treated a missing stamp as clean, so every existing checkout re-archived old 32-bit objects into a “wasm64” library. FFmpeg’s configure also compiles its probe with the compiler flags but links with only the linker flags, so the probe object was 64-bit and the link stayed 32-bit.
Fix. -m64 in both compiler and linker flags. Each tree carries a stamp, and the guard’s rule is written in the script: no stamp means unknown, not clean. Separate source and output directories per pointer model. We used -m64 instead of -sMEMORY64=1 because the former reaches CMake probe compiles.
Consequence. Before we removed the 32-bit target entirely, the first v1.2.0 release cut shipped a 32-bit engine under a client that gated on Memory64.
Lesson. Give each pointer model its own trees, and treat unknown provenance as dirty. Every artifact named on a hand-assembled link line needs a freshness check.
Story 4. Black screens with no error
Symptom. Black sky, black smoke, black interiors, no fog. Later, blank render-to-texture previews, a black minimap and a scene that rendered black under MSAA.
Why hard. There was no GL error and no shader failure. The values simply arrived as zero.
Cause. Several independent bugs with the same look:
- Struct-member uniforms such as
osg_FrontMaterial.diffuseandosg_Fog.colorsilently read as zero on WebGL2 through ANGLE. It took three commits, one per struct. - OSG unconditionally emitted
GL_NONEfor draw buffers, so every render-to-texture camera discarded all colour output. Our build notes call this the most important OSG fix. - Unsized renderbuffer formats leave a 0×0 buffer under WebGL2. Attaching one buffer to separate depth and stencil points gives
0x8CD6. - A clip-plane uniform is not per-pass state. The main pass inherited the water camera’s stale plane and discarded all geometry above the water.
- A capability flag latched true because a function existed, though the framebuffer had one attachment, so alpha-blended smoke composited as opaque black.
Fix. Flatten struct uniforms to plain ones, emit draw buffers only when a colour attachment exists, use sized formats and a single depth-stencil attachment, and publish a neutral clip plane on the root.
The minimap cost fifteen suspects and fifteen builds. The commit that closes it records the meta-lesson: the first question should have been what the feature looked like when it worked, not what changed today. A glReadPixels probe added during the hunt produced two confident, meaningless zeroes.
Lesson. On ANGLE, black usually means a silently zero uniform, an incomplete framebuffer or a dropped draw buffer. Flatten struct uniforms, check framebuffer completeness explicitly, and keep a same-binary control.
Story 5. Threading without Asyncify
We never used Asyncify. Everything runs on one GL thread, and every blocking wait became inline work, a polled state, or a spin on shared memory. Getting there took five separate deadlocks:
| Symptom | Cause | Fix |
|---|---|---|
| Intro-video skip hung about half the time | A wakeup sent without holding the queue mutex was lost, so the join never returned | Timed wait, re-checked every 10 ms. Fixed 5 of 5 runs |
| Deadlock only with a running audio context | Emscripten’s OpenAL proxies every call from a worker to the main thread. The stream thread held a mutex across those calls while main waited on it | No stream thread on the web. Streams refill inline each frame |
| Frozen on the last frame of a video | With audio inline, a blocking queue read waits forever at end of stream | Non-blocking read with three return states, padding silence on starvation |
| Equipping a weapon froze the browser | A preload worker held a mutex while main waited on it through a futex | Work queue runs inline by default. Lua also runs inline |
| Options → OK froze the game hard | The settings save ran from inside the close-event dispatch | Deferred to a clean event-loop stack |
Measurement backed the choice. With physics at 0.19 ms and Lua at 0.27 ms of a 9.14 ms frame, threading the work queue could not pay for its risk. Culling and drawing are about 65% of the frame.
We also tried the opposite direction. An experiment (OMW_PROXY=1) moves the engine to a worker with an OffscreenCanvas. Its gate questions were answered with a probe page before any porting: a blocking Atomics.wait off the main thread works, and WebGL2 on a transferred canvas works on the real GPU path. The port itself hit two walls: 24 EM_ASM blocks addressed window, which a worker lacks, and SDL2’s Emscripten backend creates its context through EGL, which is bound to the main-thread canvas. It reached a first frame on a real GPU at 23 fps, and it is off by default. We do not count it as shipped.
Lesson. Audit every blocking primitive for who can be waited on. On a proxied-GL, single-thread build the answer is often to run the work inline and add a timed re-check to any wait. Answer feasibility questions with a probe page before you port.
Story 6. Streaming four gigabytes to a thread that cannot wait
The engine reads files synchronously on the main thread. Modern Chrome forbids synchronous XHR there, and Atomics.wait is banned, so a helper worker does the fetching while the main thread spins on a shared flag (Figure 1). Every cache miss is frame time, so chunk size and cache depth had to be tuned as a pair. Our measurements, on a cold Balmora boot:
| Configuration | Misses | Main-thread stall | Note |
|---|---|---|---|
| 128-slot cache (before tuning) | 343 | 2,684 ms | 215 evictions |
| 1 MB × 384 | 259 | 2,253 ms | no evictions |
| 2 MB × 192 (shipped) | 150 | 1,380 ms | best boot |
| 4 MB × 96 | 89 | 1,346 ms | worse boot overall |
An earlier conclusion that 1 MB beat 4 MB had been measured at a cache size that thrashed, so chunk size was never the whole story. Every session also re-fetched about 300 MB because range responses do not enter the HTTP cache, so chunks now persist in the Cache API. Review found two ways for that cache to serve wrong data: a key that ignored fetch granularity gave a short read the engine treats as end of file, and re-installed mod files of identical size served stale chunks. The key now includes size, modification time and chunk size.
Lesson. In a sync-over-async design, measure stall time and evictions, tune chunk and cache size together, and key any persistent cache on everything that changes the bytes.
Story 7. Multiplayer: the page that looked alive
Symptom. A client froze permanently while its header still read “60fps (0ms/frame)”.
Why hard. The frame pump had a re-entrancy flag that was not exception-safe. One throw left it set forever, and the browser’s animation loop kept running, so the page appeared to be doing its job. A wasm trap unwound through it and a catch swallowed the trap.
Cause. A teleport gave a sound source a velocity of thousands of units per second. The Doppler clamp turned that into Infinity, and Emscripten’s OpenAL passes it straight to Web Audio, which throws The provided float value is non-finite and unwinds the whole engine frame. Desktop OpenAL simply ignores such values, which is why nobody ever noticed.
Fix. Every value handed to OpenAL goes through a finite-or check. The pump now reports a trap with a stack, shows a crash overlay, and gives the test harness an error line.
Other silent failures in the same area: the browser socket is deleted inside its own close, so onclose never fires and a reconnect watchdog hung up every frame for the rest of the session. A throwing Lua handler disables its whole subsystem, and the player then looks frozen to everyone else while their own screen looks normal. Switching worlds now reboots the page, because no list of state to undo can be trusted to be complete.
Lesson. Never swallow a trap. Put the failure path into your logging before you form a hypothesis, and read the client log for
Lua errorfirst.
Smaller lessons
- Silent audio for months. The staged FFmpeg archive predated the decoder change, and FFmpeg is the engine’s only decoder. The link script now checks for
ff_mp3_decoder. - A stub with the wrong signature becomes a trap. A mismatched X11 stub called from OSG’s static constructor traps at boot, before
main(). - A known-bad flag. Link-time optimisation miscompiles boot in our build. We have no diagnosis, only the prohibition, written in three places.
- CI memory. Ninja ran with no
-jand 34 clang jobs killed the build host while the container reported a clean exit. Onlydmesgtold the truth. Budget about 1 GB per job. - Wrong capacity numbers. We once published avatar costs an order of magnitude too high, measured while the host was at load 54 to 131. The README now carries the correction.
4. Multiplayer architecture
The server validates and relays. It does not simulate. A native, headless copy of the same engine, the sim-peer, joins as a normal client and resolves combat, falls and loot once. A modified browser client therefore cannot decide what an NPC did.
server/PROTOCOL.md and openmw/files/data/scripts/mp/.- Client A sends raw input at about 30 Hz (
PlayerInput, 12 bytes with an input sequence number) and keeps simulating locally, so input lag is zero. - The server forwards it to the sim-peer, which drives A’s body.
- The peer returns poses at about 20 Hz (
AvatarMoveBatch). The server accepts them only for players who sent input within the last two seconds. - The server returns A’s own state with the last input sequence it processed. A keeps a 64-entry ring of where it stood at each sequence number, and compares. A divergence under 256 units is corrected by 25% per frame. Beyond that it snaps, with a two-second cooldown.
- Everyone else receives A’s pose on the broadcast tick, culled by distance. Client B draws A as a puppet, rendered slightly behind real time and interpolated.
With no peer holding the world, the server drops input frames and the older client-authored path takes over, controlled by the two-second freshness window alone. Measured on an idle workstation with one browser client, 64 co-located avatars ran at 48 fps with tiered detail, against 37 fps with everyone at full detail. A 24-bot soak for 30 minutes held 24 of 24 sessions, with mean ping 4 ms and maximum 85 ms.
5. Measurements, with conditions
| Measurement | Result | Conditions and caveats |
|---|---|---|
| Browser vs native CPU per frame | 5.87 ms vs 4.89 ms (1.20×) | Windows laptop, Balmora, standing still. Native is upstream 0.51.0, the web build is 0.52.0-dev, and the web scene has a longer view distance (16384 vs 7168). The web number is the median of 12 samples of the engine’s own frame timer. |
| Where the frame goes | cull + draw ≈ 65% | Median of a settled scene, 9.14 ms total, with performance counters armed. Engine subsystems combined are 1.9 ms (21%). |
| wasm64 vs wasm32 CPU per frame | 0.785 vs 0.620 ms (+27%) | 3 runs each, non-overlapping. The commit does not record machine or scene, so treat it as indicative and re-measure before quoting externally. |
| Binary size, wasm64 vs wasm32 | 33.6 MB vs 31.3 MB (+7.3%) | Different tree states, so the delta is approximate. |
| Boot to first frame | 8.9 to 10.3 s | Balmora, retail data streamed, dev machine. About 80% of it happens after the module is up. Cold versus warm cache timing has not been recorded. |
| GL traffic | 4,784 calls, 789 draws / frame | Sampler-uniform interning cut glUniform1iv calls per draw by 85% and total GL calls by about 14%. We do not claim a frame-time gain from it. |
| Client memory | 1.5 GB heap | Retail Morrowind at Balmora. The heap never grew during a full session. Streaming cache 302 MB of a 384 MB cap. |
| Memory ceiling | crossed 4 GiB | The engine’s own allocator reached 4.50 GiB, read back a write at 4.15 GiB and kept rendering (verify-wasm64, 13 of 13 checks). |
What we have not measured. We have no recorded browser heap for Tamriel Rebuilt, the large mod that motivated wasm64, only proof that the allocator crosses 4 GiB. We have no SIMD-only comparison, no Firefox or Safari numbers, and no cold-versus-warm load time. We also reverted two changes we had counted as gains, running the frame directly in the animation callback and a light-distance setting, because both broke real keyboard input. Only one scenario presses real keys, and it caught both.
6. Where this stands against prior work
We searched before writing this, and we state what we found.
| Area | Earlier work found | What we claim |
|---|---|---|
| OpenMW in a browser | A 2023 wasm32 attempt (Pospelove/openmw-web) that never drew a frame. Two forks of our repository, neither with changes of its own. | The first publicly playable OpenMW in a browser. Our public repository begins 2026-07-02. |
| OSG on WebAssembly | Upstream OSG on Emscripten in 2017. osgVerse built OSG for wasm64 on 2024-11-14. | OSG 3.6.x running OpenMW’s shader pipeline on WebGL2 under wasm64. We do not claim the first OSG on wasm64. |
| MyGUI | MyGUI 3.4.0 (2020) added an Emscripten backend. | MyGUI on wasm64 inside a full engine. No earlier wasm64 build found. |
| Bullet | ammo.js (wasm32, from 2011). Emscripten’s Bullet port is tested under wasm64. | Bullet 3.25 in double precision for a full game. No other public build found. |
| FFmpeg, Recast, Boost | ffmpeg.wasm (2019), recast-navigation-js and a Conan recipe fix for Boost on wasm64 (October 2025), all outside a game engine. | Built for wasm64 as part of one engine stack. No earlier engine-scale example found. |
| Browser multiplayer | TES3MP is desktop-only. We found no browser client for it. | A browser client for an authoritative OpenMW-style server. None found in our search. |
Godot and Unity both offer wasm64 targets, so we do not claim wasm64 itself. We could not search the OpenMW forum, its GitLab issues, YouTube, Reddit or Hacker News directly, so “none found” means none found with our queries.
7. Limits today
- Browsers. Desktop Chrome or Edge 133 or later, or Brave. The launcher also lists Firefox 134 and later, but every measurement in this paper is from Chromium and we have no Firefox numbers yet. Safari and iOS lack Memory64.
- Hosting. Threads need cross-origin isolation headers. The page also needs WebGL2 and the File System Access API.
- Rendering. No clustered lighting, because WebGL2 has no storage buffers. GPU skinning was tried and reverted, and heads and actors are skinned on the CPU.
- Multiplayer. Labelled experimental. Verified to 48 concurrent players by the test harness, not 64. No touch controls.
8. Open source and compliance
- Licence. The engine and our adaptations are GPL-3.0-or-later. Every release ships its corresponding source.
- Dependencies. MyGUI (MIT), Bullet (zlib), Boost (BSL) and Lua (MIT) are permissively licensed. The FFmpeg build uses the minimal LGPL configuration described in section 2.
- Game data. We ship none. Players bring their own copy of Morrowind. The demo world uses openly licensed data.
9. What support goes toward
The patches and build recipes are useful to others only if they are upstreamed, documented and kept current. That is the work support goes toward:
| Work | Outcome |
|---|---|
| Upstream patches | Submit the render-to-texture, MSAA and mipmap fixes to OpenSceneGraph and the char_traits fix to MyGUI. Publish the Bullet and FFmpeg recipes as documented builds. |
| Reproducible builds | Pin every dependency, remove stale build notes, and publish the source tarball for each release with a verification script. |
| Measurements | Record wasm32 against wasm64, cold against warm load, Firefox, and a Tamriel Rebuilt heap run, all with stated conditions. |
| Browser reach | Track Safari’s Memory64 work and verify Firefox properly. |
| Rendering | Evaluate a WebGPU path, which would allow storage buffers and clustered lighting. |
These are the gaps this paper lists as open. Support goes toward the items above. It does not buy features outside them, or the removal of anyone’s contributions.
If this is useful to you, you can support the work monthly on Patreon or give once on Ko-fi. You can also read what support pays for, or ask a question in our Discord.
Sources
- Project record:
server/docs/STATUS.md,CHANGELOG.md,WASM_ADAPTATIONS.md,docs/BUILDING.md,server/README.md(Measured capacity),server/PROTOCOL.md,play/streamfs.js, and commits 609da161, a9a1a16d, 14930edc, 15c08551, e740c925, 17eb433e, 65911bfe, df7af651, 5094429e, b358c9f0, 42f6df8a, 8e50f47b, 886e2673 in github.com/Virtastic/openmw-web - Figma, “WebAssembly cut Figma’s load time by 3x” (2017-06-08), figma.com/blog/webassembly-cut-figmas-load-time-by-3x
- Autodesk, “WebAssembly at Autodesk”, QCon New York 2018
- osgVerse Memory64 commit b90f33e0 (2024-11-14), github.com/xarray/osgverse
- OpenSceneGraph osgemscripten example (2017-06-22); MyGUI 3.4.0 release notes (2020-02-11)
- ammo.js (github.com/kripken/ammo.js); ffmpeg.wasm (github.com/ffmpegwasm/ffmpeg.wasm); recast-navigation-js
- conan-center-index PR 28655, Boost on wasm64 (2025-10-22)
- Pospelove/openmw-web (2023); Godot PR 102378 (2026-02-26); Unity 6.6 manual, “WebAssembly 64-bit support”
- Browser support: caniuse.com/wf-wasm-memory64
Have a desktop app that belongs in the browser?
We recompile legacy desktop software to WebAssembly, then maintain and host it. See it proven on your own app.