WASM at Edge

Architecture

Compiling application code to WebAssembly and running it on CDN edge servers as a short, sandboxed, memory-capped request handler. Startup is near-instant because no VM or container boots per request. Fastly Compute runs Wasm natively; Cloudflare Workers runs it inside a V8 isolate.

Also known as WebAssembly at the edge.

10 min read Updated Aug 30, 2026

Full Explanation

WebAssembly at the edge (Wasm for short) means compiling your application code to WebAssembly. You then run it on a CDN's edge servers instead of in a centralised cloud region. WebAssembly itself is a binary instruction format for a stack-based virtual machine. It is designed as a portable compilation target for programming languages, so one binary runs on any host that implements the standard. That portability, plus a memory-safe, sandboxed execution environment, is what lets a CDN accept code written in Rust, Go, JavaScript or C++. The CDN can run that code safely on shared machines throughout its network of points of presence.

It is one way to build an edge function, not a synonym for it. Cloudflare Workers is a JavaScript isolate with Wasm as an option inside it, and Akamai's edge functions product, EdgeWorkers, runs JavaScript. It is also not a general compute platform. Budgets are tight, there are no threads, and nothing on local disk outlives the request. So treat it as a very fast, very short-lived request handler, and keep long-running or stateful work somewhere else.

How it works

The edge platform embeds a WebAssembly runtime. Your code is just a module that runtime loads. The platform decides when a sandbox starts and how long it lives.

  • Write a request handler in a language that compiles to Wasm. Fastly recommends its official SDKs, available for Rust, JavaScript, Go and C++. Cloudflare documents compiling languages such as Rust, Go or C to a module you call from a Worker.
  • Compile and ship it. On Fastly, fastly compute build runs your language toolchain and creates the package. fastly compute deploy uploads it and attaches it to a service.
  • Fastly Compute runs the module under the Wasmtime runtime. In the default configuration, each request starts a new sandbox. This gives fast startup and per-request isolation.
  • Cloudflare does it differently. A Worker runs in a V8 isolate: a lightweight context created inside an already-running process, not a fresh virtual machine. This is what eliminates VM-style cold starts. Your Wasm module is instantiated inside that isolate. You can do this with WebAssembly.instantiate(), or by writing the whole Worker in Rust against bindings that expose the Workers JavaScript APIs.
  • The module has no ambient authority: it can only reach what the host grants it. Outbound HTTP goes through the platform's own path. On Fastly, that path is a backend configured on the service, or one registered at runtime with Dynamic Backends. On Cloudflare, it is fetch(), which counts as a subrequest.
  • The handler returns a response and the run ends. Anything that must outlive the request goes to a platform data store, not to memory or disk.

The portable version of "what the host grants" is the WebAssembly System Interface (WASI). WASI is a group of standards-track API specifications covering networking, filesystem access and more. In WASI, a module starts with no ambient authority and can only do what the host explicitly allows. WASI is not yet how you write edge code in practice. It has reached three milestone releases: 0.1, 0.2 and 0.3. On Cloudflare Workers, WASI support is experimental, with only some syscalls implemented. Day to day, you use the vendor's SDK.

Why it matters for a CDN

A CDN answers requests in single-digit milliseconds. So per-request compute is only viable if starting it is nearly free. That is the entire argument for Wasm at the edge. The sandbox or isolate is created inside a process that is already running, so nothing boots a virtual machine or a container per request. Fastly describes its isolation technology as creating and destroying a sandbox for each request in microseconds, and markets "zero millisecond cold starts". Cloudflare says its isolate model "eliminates the cold starts of the virtual machine model". Akamai claims sub-millisecond Wasm cold start. Those are vendor figures. But the architectural point under them is real. It is why running code on every request at every location is affordable at all.

Second, handling the request at the edge removes a round trip to the origin for any work that does not need the origin. Fastly's own summary of the workload includes authentication, personalisation, geofencing, SEO, observability, templating, APIs, "or even your entire application". That list maps well to what people actually deploy. Examples include checking a JWT before a request is allowed anywhere near the origin, geolocation-driven routing and personalisation, A/B bucketing, and composing several backend calls into one response as an API gateway layer.

What CDNs do

  • Fastly Compute is Wasm-native. Fastly states that security and portability come from compiling your code to WebAssembly, and that it runs your code using Wasmtime. Every request gets a new sandbox by default. An opt-in reusable-sandbox mode lets one sandbox serve several requests, to avoid repeating expensive initialisation. Fastly's default limits per execution are 50 ms of CPU, 1M bytes of stack and a 128 MB heap. Other per-execution defaults are 2 minutes of runtime (60 s on trial accounts) and 32 backend requests (10 on trial), with compiled packages capped at 100 MB per service. Those are defaults that change with account type, and can be raised on request.
  • Cloudflare Workers offers Wasm inside its V8 isolates, rather than as the primary runtime. SIMD is supported, threading is not, and WASI support is experimental. Memory is 128 MB per isolate, including the JavaScript heap and Wasm allocations. That limit is per isolate, not per invocation, and one isolate can serve many concurrent requests. CPU time is 10 ms per request on the Free plan, and up to 5 minutes on Paid (30 s by default). Compressed Worker size is capped at 3 MB on Free and 10 MB on Paid, and startup time is capped at 1 second.
  • Akamai separates the two. EdgeWorkers, its edge-server functions product, runs JavaScript, with cold start documented as currently under five milliseconds. Its WebAssembly offering is a managed serverless platform on Akamai Cloud, powered by Fermyon WebAssembly Functions and built on the open-source Spin toolchain. Access is arranged through an Akamai account team. Akamai announced its acquisition of Fermyon in December 2025.

Watch out for

  • Per-request isolation is a Fastly default, not a law. Turn on reusable sandboxes, and, in Fastly's words, per-request isolation is not guaranteed. Sandbox state such as global variables may persist across requests. Never keep a password or token in a global. Never lazily initialise a global from one request's data.
  • On Cloudflare, sharing is the default. An isolate is not per-request, so mutable globals are visible to later requests in the same isolate. Cloudflare advises against storing mutable state in global scope, unless you have accounted for the isolate being evicted. Eviction can happen because of resource limits on the machine, a script that looks like it is probing the sandbox, or the Worker's own resource limits. This is the same hazard as Fastly's opt-in mode, with the opposite default.
  • Read every limit's scope, not just its number. Fastly's 50 ms is CPU time per execution, not wall clock. The same handler may sit for up to 2 minutes waiting on backends. Cloudflare's CPU limit likewise excludes time spent waiting on fetch(), KV reads or database queries, and its 128 MB is per isolate rather than per request. Figures also differ by plan and account type. So a limit quoted without its plan is not something you can design against.
  • Wasm binaries cost startup time. Compiling to WebAssembly often pulls in extra runtime dependencies. So Workers that use Wasm are typically larger than an equivalent JavaScript Worker. The larger the Worker, the longer it may take to start, against a 1-second startup budget and a 3 MB compressed size cap on the Free plan. Cloudflare recommends tools such as wasm-opt to shrink the binary.
  • No threads. Threading is not possible in Workers: each Worker runs in a single thread, and the Web Worker API is not supported. On Fastly, a sandbox is never reused while it is handling a request, even when that request is idle waiting on asynchronous work. So a sandbox never processes two requests concurrently. Design one request-to-response unit, not a worker pool.
  • Portable does not mean portable today. Workers generally supports the same WebAssembly feature set as Google Chrome. WASI support there is experimental. The core standard is still moving too: Wasm 3.0 became the live standard in September 2025, adding a 64-bit address space, multiple memories and garbage collection. Expect to adjust a module when you move it between vendors.
  • Check whether the platform already ships the feature. Image transformation is the use case everyone attaches to edge Wasm. On Fastly, though, it is a separate product with its own documented Compute-platform limitations, and image optimization is usually cheaper bought than written.

Best practice

  • Keep the handler short and stateless. Put anything that must outlive a request in a platform store: Fastly Config Store, KV Store or Secret Store, or Cloudflare KV, R2 or D1. Stream request and response bodies rather than buffering whole payloads in memory.
  • Choose Wasm for a language reason. If the logic is JavaScript-shaped, the platform's JavaScript path is smaller and starts faster. Wasm earns its place when you want Rust, Go or C++, or an existing library you are not going to rewrite.
  • Measure and shrink the binary before you ship it. On Cloudflare, wrangler deploy --outdir bundled/ --dry-run reports the compressed size. Run wasm-opt on the module. Size feeds directly into startup time and into the size caps.
  • Load-test against the limits for the plan you are actually on. Check whether a limit you are close to can be raised, before you design around it.
  • If you enable Fastly's reusable sandboxes, treat every global as shared. Initialise globals only from public data: a country-code lookup is safe, a user's session is not. Build dynamic backend names deterministically from their configuration. Keep per-request secrets out of anything that survives the request. Fastly also warns about timing side channels. In a timing side channel, one request's execution time lets the next request in the same sandbox infer something about it.

Examples

# Fastly Compute: Rust edge handler
use fastly::{Error, Request, Response};

#[fastly::main]
fn main(req: Request) -> Result<Response, Error> {
    // Geolocation at the edge
    let geo = req.get_client_ip_addr()
        .and_then(|ip| fastly::geo::geo_lookup(ip));

    let country = geo.map(|g| g.country_code())
        .unwrap_or("US");

    // Route to nearest backend
    let backend = match country {
        "JP" | "KR" | "SG" => "origin-apac",
        "DE" | "FR" | "GB" => "origin-eu",
        _ => "origin-us",
    };

    let mut beresp = req.send(backend)?;
    beresp.set_header("X-Edge-Country", country);
    Ok(beresp)
}

# Cloudflare Worker with WASM (Rust via wasm-bindgen)
use worker::*;

#[event(fetch)]
async fn main(req: Request, env: Env, _ctx: Context)
  -> Result<Response> {
    let url = req.url()?;
    if url.path().starts_with("/api/") {
        // Transform API response at edge
        let mut resp = Fetch::Url(url).send().await?;
        let body: serde_json::Value = resp.json().await?;
        let filtered = filter_fields(&body);
        Response::from_json(&filtered)
    } else {
        Response::error("Not found", 404)
    }
}

# Deploy WASM to Fastly
fastly compute init --language rust
fastly compute build
fastly compute deploy

# Deploy to Cloudflare
npx wrangler deploy

Frequently Asked Questions

Compiling application code to WebAssembly and running it on CDN edge servers as a short, sandboxed, memory-capped request handler. Startup is near-instant because no VM or container boots per request. Fastly Compute runs Wasm natively; Cloudflare Workers runs it inside a V8 isolate.

# Fastly Compute: Rust edge handler
use fastly::{Error, Request, Response};

#[fastly::main]
fn main(req: Request) -> Result<Response, Error> {
    // Geolocation at the edge
    let geo = req.get_client_ip_addr()
        .and_then(|ip| fastly::geo::geo_lookup(ip));

    let country = geo.map(|g| g.country_code())
        .unwrap_or("US");

    // Route to nearest backend
    let backend = match country {
        "JP" | "KR" | "SG" => "origin-apac",
        "DE" | "FR" | "GB" => "origin-eu",
        _ => "origin-us",
    };

    let mut beresp = req.send(backend)?;
    beresp.set_header("X-Edge-Country", country);
    Ok(beresp)
}

# Cloudflare Worker with WASM (Rust via wasm-bindgen)
use worker::*;

#[event(fetch)]
async fn main(req: Request, env: Env, _ctx: Context)
  -> Result<Response> {
    let url = req.url()?;
    if url.path().starts_with("/api/") {
        // Transform API response at edge
        let mut resp = Fetch::Url(url).send().await?;
        let body: serde_json::Value = resp.json().await?;
        let filtered = filter_fields(&body);
        Response::from_json(&filtered)
    } else {
        Response::error("Not found", 404)
    }
}

# Deploy WASM to Fastly
fastly compute init --language rust
fastly compute build
fastly compute deploy

# Deploy to Cloudflare
npx wrangler deploy

Yes. WASM at Edge is also known as WebAssembly at the edge. Compiling application code to WebAssembly and running it on CDN edge servers as a short, sandboxed, memory-capped request handler. Startup is near-instant because no VM or container boots per request. Fastly Compute runs Wasm natively; Cloudflare Workers runs it inside a V8 isolate.

Related CDN concepts include:

  • API Gateway — An API gateway is the single front door for API traffic: a reverse proxy that …