Skip to main content

Chapter 11: Realtime: Audio and Video at the Edge

How do I build applications with real-time audio and video?


A shared document needs timely state updates; a video call needs a continuous stream of playable media. Both are real-time applications, but they tolerate different failures. That distinction determines the transport and the infrastructure you need.

Durable Objects with WebSockets can coordinate the shared document and its connected clients, as Chapter 7 explains. But audio and video streaming needs a different stack: the protocols differ (WebRTC rather than WebSockets), the infrastructure differs (media servers rather than application servers), and the operational complexity differs (NAT traversal, codec negotiation, adaptive bitrate). Cloudflare Realtime exists because Durable Objects are not designed for media.

WebSockets carry reliable, ordered messages while a connection remains healthy. Interactive media also needs congestion control, loss recovery and adaptive quality. WebRTC supplies those mechanisms; application state still needs its own acknowledgement and recovery rules.

This chapter covers when you need Cloudflare Realtime versus when Durable Objects suffice, explains why anycast selective forwarding units (SFUs) change the trade-offs, describes the operational realities of running real-time media in production, and shows how to evaluate whether Cloudflare fits your requirements.

Two kinds of real-time​

Understanding the distinction between data real-time and media real-time shapes every architectural decision you make.

Data real-time involves application messages and shared state. The acceptable delay depends on the operation: a cursor, chat message and trading decision have different budgets. WebSockets carry messages; the application still defines acknowledgement, reconciliation and what to do after a disconnect.

Media real-time involves continuous audio and video. The latency budget includes capture, encoding, network transit, buffering and playback. Delays that are harmless in a dashboard interrupt turn-taking in a conversation. WebRTC provides media transport and adaptation; SFUs forward streams, and TURN relays help clients cross restrictive networks.

The protocol reality​

TCP delivers WebSocket bytes in order. A lost packet can therefore delay later data while it is retransmitted. WebRTC media transport can recover or conceal some loss while prioritising data that is still useful for playback. How well this works depends on the codec, congestion control and the pattern of packet loss.

Use WebSockets for messages whose ordering your application needs. Use a media stack for interactive audio and video, where adapting bitrate and deciding which late packets to discard are part of correctness.

When Durable Objects suffice​

Many "real-time" applications don't actually need media infrastructure, so before reaching for Cloudflare Realtime, you should confirm that Durable Objects with WebSockets can't solve your problem.

Text chat and messaging: WebSockets through Durable Objects handle this perfectly. Messages are small, latency tolerance is generous, and the coordination model (one Durable Object per conversation) maps naturally to the domain. Cloudflare Realtime adds cost and complexity with no benefit.

Collaborative editing: Whether you're synchronising document state, whiteboard drawings, or design tool operations, WebSockets suffice. The data volumes are modest, sub-second latency is acceptable, and the interesting problems (conflict resolution, operational transformation) live in your application logic, not the transport layer.

Live dashboards and data feeds: Stock tickers, sports scores, IoT sensor readings, and analytics dashboards involve one-way data flow from server to many clients. WebSockets handle fan-out efficiently. Durable Objects coordinate the source of truth.

Multiplayer game state: Turn-based games, strategy games, and even many action games synchronise state through WebSockets. The game logic determines what "real-time" means; unless you're building a first-person shooter requiring sub-50ms latency, Durable Objects likely suffice.

Presence and cursor tracking: Showing who's online and where their cursor is involves small, frequent updates. WebSockets handle this trivially. The coordination challenge (ensuring all participants see consistent presence) is exactly what Durable Objects solve.

For these cases, Chapter 7’s object-per-entity pattern provides one owner for a conversation, document or game session, with WebSockets connecting participants. The application still defines acknowledgements, conflict handling and recovery, and protects transitions across external awaits. Choose this when that ownership model fits the interaction.

When you need Cloudflare Realtime​

Cloudflare Realtime becomes necessary when your application involves actual audio or video between participants rather than just data synchronisation.

Video conferencing: An SFU lets a participant publish a stream once and forwards it to subscribers. In a full mesh of N participants, each sender needs N − 1 outbound streams. That growing upload burden, and the client’s decoding capacity, make an SFU useful well before a fixed participant limit can describe every call.

Voice calling: Audio uses less bandwidth than video, but conversational turn-taking still depends on low delay and stable playback. Measure the path under packet loss and jitter, not just its average bitrate.

Live streaming with interaction: A presenter broadcasting to an audience, with the audience able to respond via voice or video. The broadcast might tolerate higher latency (1-2 seconds is common for live streams), but interactive segments need the same low latency as a call.

AI voice agents: The delay includes endpointing, speech recognition, model response and speech synthesis. Test interruptions and turn-taking as well as response time. A fast media path cannot compensate for a pipeline that waits too long to decide the user has finished speaking.

Telehealth and remote assistance: Video calls with specific requirements around quality, reliability, and sometimes recording. Medical consultations, technical support with screen sharing, and remote expert guidance all need robust video infrastructure.

For these use cases, Durable Objects with WebSockets cannot help. You need WebRTC infrastructure: media servers, TURN relays, and client SDKs that handle codec negotiation, adaptive bitrate, and network condition changes.

The anycast SFU architecture​

A regional SFU makes media placement an application decision: which location should serve a call, and what happens when participants are far from it? Managed services such as Chime SDK operate the media capacity for you; self-hosted SFUs also leave scaling and failover to your team.

Cloudflare Realtime routes participants into its distributed SFU infrastructure using anycast. Your application does not select a media region or provision SFU servers. This simplifies deployment, but actual paths still depend on network routing and the locations of the people in the call.

The traditional SFU problem​

In a traditional SFU deployment, all participants in a call connect to the same server. If three people in London and one person in Sydney join a call, the SFU must be somewhere. Place it in London, and the Sydney participant has 300ms round-trip latency. Place it in Sydney, and the London participants suffer. Place it somewhere in between, and everyone has mediocre latency.

Some systems try to solve this by selecting SFU location dynamically based on the first participant or the geographic centre of all participants. This helps when participants cluster geographically, but creates suboptimal routing when they don't, and "geographic centre" between London and Sydney is somewhere in the Indian Ocean where no data centres exist.

Cloudflare's approach​

Cloudflare lets each participant enter the media network near their own connection, then forwards media between locations. London participants need not send their local exchange through the same distant regional server chosen for a Sydney participant.

The London-to-Sydney conversation still crosses continents. Distributed entry points can avoid unnecessary detours; they cannot remove propagation delay, congestion or a participant’s poor last-mile connection.

Distributed entry points can reduce regional detours

Loading diagram…

Cloudflare manages SFU capacity and distribution. Your application still controls the subscription graph: how many streams each participant receives, at what quality, and when inactive video can stop. More recipients mean more forwarded media even when server placement is automatic.

Cascading for large calls​

For calls with many participants, Cloudflare uses cascading: media flows through a tree of servers rather than a single hub. Participants in London connect to London servers; those servers cascade to servers handling other regions. This keeps latency low while enabling calls with hundreds of participants.

The cascading topology forms automatically based on participant locations. You don't configure it; you don't even see it. From your application's perspective, participants join a call and can see and hear each other. The platform handles routing optimisation invisibly.

Latency characteristics​

Measure media latency on the paths your users take. Network round-trip time is only one component; capture, encoding, jitter buffering and decoding contribute to what participants experience. Measure connection success, time to first media and sustained quality separately.

A benchmark should state the participants’ locations, devices, codecs and network conditions. A single regional average cannot tell you whether a cross-continent call or a user behind a corporate firewall will work.

TURN: solving NAT traversal​

WebRTC wants peer-to-peer connections. In practice, NATs and firewalls often prevent direct connectivity. TURN (Traversal Using Relays around NAT) servers act as intermediaries, relaying traffic when direct connection fails.

TURN is essential for users whose network blocks a direct media path. Test corporate networks and mobile connections explicitly, and measure relay usage in your own audience; a consumer traffic mix is a poor capacity estimate for an enterprise calling product.

Cloudflare provides TURN as part of Realtime. The same anycast architecture applies: TURN allocations route to the nearest Cloudflare location. The service handles the complexity of maintaining relay allocations, managing credentials, and ensuring connectivity.

STUN vs TURN​

STUN (Session Traversal Utilities for NAT) helps peers discover their public IP addresses and attempt direct connection. It's lightweight and free on Cloudflare (stun.cloudflare.com). Most connections succeed with STUN alone: the peers discover their public addresses, exchange candidates through signalling, and establish direct media flow.

TURN relays actual media traffic when STUN-facilitated direct connection fails. This happens when both peers are behind symmetric NATs, when firewall rules block UDP, or when network policies prevent direct peer connections. TURN consumes bandwidth because it relays all media, and it costs money ($0.05 per GB).

The fallback sequence: try direct connection first, fall back to TURN only when necessary. Your application doesn't control this; the WebRTC stack handles it automatically. But understanding the sequence helps when debugging connectivity issues and predicting costs.

A TURN relay forwards encrypted WebRTC traffic between its endpoints without needing the media keys. That is a property of the relay path, not proof that an SFU-based meeting is end-to-end encrypted between participants. Establish where media encryption terminates for the architecture you deploy.

Estimating TURN usage​

Measure the proportion of sessions and bytes using TURN in a representative pilot. Restrictive enterprise networks can produce a different relay mix from home connections. Include failed joins in the measurement: a low relay share is not reassuring if those users never connected.

RealtimeKit: the client SDK​

RealtimeKit handles client media negotiation, track management and common meeting behaviour. That saves integration work, but your application still owns identity, meeting permissions, recovery and the user experience around failures. Its participant-minute rates are covered below.

The SDK is available for web (React, plain JavaScript, and Angular), React Native, Flutter, iOS, and Android. It provides two layers:

Core SDK: APIs for joining sessions, managing media tracks, handling events. You build your own UI but don't implement WebRTC negotiation, ICE candidate handling, or codec selection.

UI Kit: Pre-built components for common patterns. Video grids, participant lists, control bars. Useful for rapid prototyping or applications where custom UI isn't a differentiator.

Basic integration​

The integration model splits cleanly between server and client. Your backend creates a meeting and adds participants through the REST API; adding a participant returns an auth token scoped to that person and their role. The client initialises the SDK with that token and joins. Account credentials never reach the client. Here addParticipant is an application helper that calls the REST API and returns its participant data. The backend must authenticate the caller and authorise the requested meeting and preset before issuing the token.

Server: return the authorised participant’s token
const participant = await addParticipant(meetingId, {
name: "Jamie",
custom_participant_id: authenticatedUser.id,
preset_name: "group_call_host"
});
return Response.json({ authToken: participant.token });

The browser retrieves that token from its authenticated application endpoint. The following code runs in the browser, separately from the server handler:

Browser: use the participant token to join
import RealtimeKitClient from "@cloudflare/realtimekit";

const response = await fetch("/api/meeting-token", { method: "POST" });
if (!response.ok) throw new Error("Unable to obtain a meeting token");
const { authToken } = await response.json();
const meeting = await RealtimeKitClient.init({ authToken });
await meeting.join();

The preset named when adding a participant deserves more attention than it usually gets. Presets define what a participant can do (publish video, share a screen, kick others) and whether they count as audio-only or audio-and-video, which is also how they're priced. Preset design is simultaneously your permission model and your billing model.

Network resilience and media failures​

Real-world networks are unreliable: connections drop, quality degrades, and users switch between WiFi and cellular mid-call. RealtimeKit reconnects automatically through transient failures and exposes connection state so your interface can tell users what's happening; persistent failures eventually surface as a terminal failed state that your application should meet with a retry offer or a graceful exit. You write the messaging, not the recovery logic.

The SDK adapts media to available bandwidth, and simulcast can let receivers subscribe to different quality layers. Additional layers consume sender bandwidth; measure their cost rather than assuming a fixed uplift. Decide when the application should reduce outgoing quality or disable video.

Media applications additionally fail in ways web applications don't: camera permission denied, microphone already in use by another application, Bluetooth devices disconnecting silently. The SDK surfaces these as distinct errors; handle each with a specific message, because a generic "unable to access camera" helps nobody whose actual problem is that another application holds the device. And always verify that a device switch took effect before telling the user it did.

Recording​

Recording video calls typically requires running a headless browser or custom media pipeline to capture streams. RealtimeKit includes recording as a platform capability: start and stop it through the REST API, and the pipeline runs on Cloudflare's infrastructure. Recordings can be composited into a single file with all participants' audio and video, captured as separate audio tracks per participant for post-production flexibility, or exported as raw RTP directly into R2 for custom processing.

This matters for compliance-driven applications (financial services requiring call recording), education (lecture capture), and content creation (podcast recording). What would otherwise require significant infrastructure becomes a configuration option.

Cost model and estimation​

Real-time media pricing varies significantly across providers. Understanding the cost model prevents surprises.

Two pricing layers​

Cloudflare Realtime prices its two layers differently, and knowing which layer you're consuming determines your cost model.

The raw SFU and TURN services charge for bandwidth: $0.05 per GB of egress, with the first 1,000 GB each month free. The allowance is shared across SFU and TURN, and traffic relayed through TURN into the SFU isn't double-charged. Bandwidth pricing creates different economics from per-minute pricing: audio-only streams at roughly 100 Kbps per participant cost dramatically less than video at 1-4 Mbps, so applications that adapt quality to conditions optimise their own bill.

RealtimeKit charges per participant-minute: $0.002 for audio and video, or $0.0005 for audio-only. Recording export costs $0.010 per minute, $0.003 for audio-only, or $0.0005 for raw RTP into R2; R2 storage is separate. The participant’s preset determines the billing category, so preset design directly shapes the bill.

Estimating bandwidth​

For a typical video call:

QualityBitrateGB per hour (one direction)
Audio only50–100 Kbps0.0225–0.045 GB
360p video400-600 Kbps0.18-0.27 GB
720p video1.5-2.5 Mbps0.68-1.13 GB
1080p video3-5 Mbps1.35-2.25 GB

For a four-person call lasting one hour, assume each participant receives three streams, each averaging 2 Mbps. The SFU sends twelve streams in total: 10.8 GB, or $0.54 before the monthly free allowance. Published media entering the SFU is free. RealtimeKit charges four participants for 60 minutes at $0.002, or $0.48 before recording. Actual SFU usage changes with subscriptions, simulcast layers and bitrate adaptation.

For raw SFU estimates, count received streams, not just participants. With everyone receiving everyone else, a call with N participants has N × (N − 1) outgoing streams. A webinar with one active publisher has a different cost curve. Convert total outgoing bits to billed GB, subtract the shared monthly allowance, then apply the egress rate.

Comparison with alternatives​

ProviderPricing Model4-person, 1-hour call
Cloudflare Realtime SFU$0.05/GB egress$0.54 at the 2 Mbps assumption above
Cloudflare RealtimeKit$0.002/participant-minute$0.48
AWS Chime SDK$0.0017/participant-minute$0.41
Azure Communication Services$0.004/participant-minute$0.96
Twilio Video$0.004/participant-minute (group)$0.96

On the raw SFU, Cloudflare's cost varies with video quality where the others charge flat per-minute rates regardless. For audio, Cloudflare is dramatically cheaper on either layer: bandwidth pricing rewards low-bitrate streams, and RealtimeKit's audio-only rate of $0.0005 per participant-minute undercuts Chime's $0.0017 and ACS's $0.004. For high-quality video, costs converge.

Hidden Costs

Factor in:

  • TURN relay: account for the actual relay path; Cloudflare does not double-charge traffic between its TURN and SFU services
  • Recording: Storage (R2 pricing) plus processing bandwidth
  • Transcoding: If you're re-encoding recordings, compute costs apply
  • Development time: Newer SDKs may require more integration effort than mature alternatives

Failure modes worth naming​

Real-time media systems fail in ways that differ from typical web applications. Naming these failures creates vocabulary for debugging and post-incident analysis.

Connectivity failures​

ICE negotiation timeout occurs when peers cannot establish a connection within the negotiation window (typically 30 seconds). Common causes: both peers behind symmetric NATs with TURN unavailable, firewall blocking UDP entirely, or incorrect TURN credentials. Symptoms: connectionState never reaches connected. The fix depends on the cause; ensure TURN is available and correctly configured.

TURN allocation exhaustion manifests when your TURN server runs out of relay allocations. Each concurrent connection through TURN consumes an allocation. Cloudflare's managed TURN scales automatically, but if you're running your own TURN infrastructure, monitor allocation counts. Symptoms: new connections fail while existing connections continue working.

Firewall state timeout affects long-running calls. Stateful firewalls track UDP sessions and drop state after inactivity (often 30 seconds). If media pauses (everyone mutes), the firewall may drop the session, requiring ICE restart. Symptoms: call drops after period of silence. Keep-alive packets mitigate this; RealtimeKit handles it automatically.

Quality degradation​

Bandwidth collapse occurs when available bandwidth drops below minimum viable quality. The codec can't reduce bitrate further without becoming unwatchable. Symptoms: extreme pixelation, then freezing, then disconnection. The only fix is more bandwidth or disabling video.

Jitter buffer overflow happens when network jitter exceeds the buffer's capacity to smooth it. Packets arrive too irregularly to maintain smooth playback. Symptoms: choppy audio, stuttering video. Increasing buffer size helps but adds latency; the trade-off is application-specific.

Asymmetric quality manifests when one participant has good upload but poor download (or vice versa). They can send clearly but receive poorly. Symptoms: "I can see you fine but you're frozen on my screen." Debugging requires checking both directions independently.

Application-level failures​

Permission persistence problems occur when browsers revoke media permissions unexpectedly. Safari on iOS revokes permissions when the page backgrounds. Chrome may prompt again after browser updates. Symptoms: camera or microphone stops working mid-call. Handle the permissiondenied error gracefully and prompt users to re-grant.

Device switching failure happens when a user selects a new camera or microphone and the switch fails silently. Common with Bluetooth devices that disconnect unexpectedly. Symptoms: user thinks they switched devices but audio still comes from the old source. Always verify device changes took effect.

Session state desynchronisation occurs when your application's state diverges from the actual session state. Participants your app thinks are present have actually left; tracks your app thinks are active have ended. Symptoms: ghost participants, stale video. Reconcile state on reconnection and handle all participant/track events.

The debugging playbook​

  1. Check connectivity first. Can both peers reach the signalling server? Can they reach TURN? Are ICE candidates being exchanged? webrtc-internals in Chrome shows all of this.

  2. Examine network stats. RealtimeKit exposes RTCPeerConnection stats. Look for packet loss, jitter, and round-trip time. High packet loss (>5%) indicates network problems. High jitter indicates unstable network. High RTT indicates geographic distance or congested paths.

  3. Verify media tracks. Are tracks actually producing frames? A muted track still exists but produces silence. A failed track produces nothing. Check track.readyState and monitor for the ended event.

  4. Test in isolation. Connect two browsers on the same machine. If that fails, the problem is configuration. If that works, add network complexity incrementally until you reproduce the failure.

Operational considerations​

Running real-time media in production requires monitoring and practices beyond typical web application operations.

Monitoring quality​

Quality problems don't always surface as errors. A call might "work" but sound terrible. Monitor:

Media quality metrics: RTCPeerConnection provides statistics on packets sent/received, packets lost, jitter, and round-trip time. Export these to your observability platform. Alert when packet loss exceeds 5% or jitter exceeds 100ms.

Call quality score: Derive a composite score from network metrics. Many applications use Mean Opinion Score (MOS) estimation based on packet loss and latency. A MOS below 3.5 indicates noticeably degraded quality.

Connection success rate: Track what percentage of join attempts result in working media. A success rate below 95% warrants investigation.

Time to first frame: How long from clicking "join" until the user sees video? Target under 3 seconds. Longer delays suggest connection issues or slow backend authentication.

The mechanics are unglamorous: sample RTCPeerConnection stats on an interval, export packet loss, jitter, and round-trip time as gauges to your observability platform, and derive your quality score from those. The decision that matters is committing to export these metrics at all, because quality problems invisible to your dashboards get diagnosed one support ticket at a time.

Capacity planning​

Cloudflare handles SFU scaling automatically, but your backend systems need capacity planning:

Session management: Each active session requires state. A Durable Object per session scales naturally, but estimate your concurrent session count for database and cache sizing.

Authentication tokens: Token generation happens on every join. If you're calling external auth services, ensure they handle peak load.

Recording storage: If recording is enabled, estimate storage growth. A one-hour 720p recording is roughly 1-2 GB. A busy platform generating thousands of hours of recordings monthly needs terabytes of storage and a retention policy.

Post-processing pipelines: If you're transcribing, analysing, or transforming recordings, those workloads scale with call volume. Queue-based processing (Chapter 9) prevents overload during peak times.

Testing real-time applications​

Real-time media quality depends heavily on network conditions. Testing on a developer's fast, stable connection misses problems users will encounter.

Network condition simulation: Test controlled packet loss, latency, jitter and bandwidth limits, including abrupt changes during a call. Choose profiles from observed user networks; one fixed “normal mobile” profile misses restrictive and unstable connections.

Geographic testing: Use actual connections from different regions. A VPN or cloud VM in distant locations reveals latency issues local testing misses.

Scale testing: Join many participants to a single call. Quality often degrades non-linearly; a call working fine with 5 participants may struggle at 15.

Device testing: Test on actual mobile devices, not just browser simulators. iOS Safari and Android Chrome have different WebRTC implementations with different quirks.

A complete example: video consultation platform​

Bringing the pieces together: a telehealth platform supporting video consultations between patients and doctors, with recording for medical records and AI-assisted transcription.

Architecture overview​

Session control and media take different paths

Loading diagram…

Session coordination​

The Durable Object manages consultation lifecycle, participant permissions, and integration with external systems. One object per consultation gives every request worldwide a single coordination point: patient and doctor authenticate against it, it creates the RealtimeKit meeting and issues role-scoped participant tokens, and it holds session state that survives disconnections.

Authorise the consultation and role before issuing a participant token. The preset should grant only the media capabilities needed by that role. Persist the consultation state separately from the media session so reconnecting clients can recover their place.

Ending the consultation and completing its recording are separate events. Mark the call ended, then wait for confirmation that the recording is available before scheduling transcription. Use the recording identifier as an idempotency key so duplicate completion events cannot create duplicate records.

Post-processing pipeline​

Once the recording is available, a queue consumer retrieves it from R2, transcribes it, and produces a draft clinical summary. Store the transcript and summary alongside the consultation record, with human review before a generated summary becomes part of the authoritative clinical record. Make each stage safe to retry (Chapter 9).

Process a recording only after it is ready

Loading diagram…

Use post-processing transcription when a complete recording and review are more valuable than immediacy. Use live transcription for captions or an agent that must react during speech. RealtimeKit can route transcription through Workers AI; evaluate supported languages, noisy audio and latency on the selected transcription model. Neither path guarantees an accurate clinical record without review.

Comparing to hyperscaler alternatives​

Understanding what you're choosing between clarifies the decision.

Compare a concrete call design across providers: participant roles, subscriptions, recording, phone connectivity, supported client devices and operational evidence. A count of network locations is not a latency benchmark, and a participant cap is meaningful only with the required number of published and received tracks. Test those limits against the intended meeting shape.

When hyperscalers win​

PSTN requirements: If users need to dial in via phone number, AWS Chime and Azure Communication Services offer integrated PSTN. Cloudflare Realtime does not; you'd need Twilio or similar for phone connectivity.

Existing AWS/Azure investment: If you're deeply integrated with AWS or Azure, the native video services integrate more smoothly with your existing authentication, storage, and monitoring.

Enterprise compliance: Large enterprise deployments often require specific certifications (FedRAMP, HIPAA BAA). Verify Cloudflare's current certification status against your requirements. The hyperscalers have longer compliance track records.

Predictable budgeting: RealtimeKit and the managed alternatives use participant-minute pricing. Raw SFU costs depend on outgoing streams and their bitrate, so estimate from representative calls rather than treating a participant-minute as a fixed amount of bandwidth.

When Cloudflare wins​

Globally distributed users: Distributed SFU entry points can reduce routing detours. Test the improvement with the locations and networks your audience actually uses.

Variable quality needs: If calls range from audio-only to high-definition video, the SFU's bandwidth pricing rewards efficiency where hyperscaler per-minute rates don't distinguish, and RealtimeKit's rates price audio-only participants at a quarter of audio-and-video ones.

Integrated platform: Workers, Durable Objects, R2 and Workers AI provide nearby services for authorisation, coordination, storage and post-processing. Realtime still needs application-specific participant credentials and permission checks.

Cost at scale for audio: Audio-heavy applications (voice agents, conferencing without video) cost dramatically less on Cloudflare: $0.0005 per participant-minute on RealtimeKit against Chime's $0.0017 and ACS's $0.004, and cheaper still on the raw SFU's bandwidth pricing.

Making the decision​

The decision framework for real-time features:

QuestionIf YesIf No
Does the application involve audio or video between participants?Evaluate RealtimeUse Durable Objects + WebSockets
Do users need to dial in via phone?Use Chime/Azure/TwilioCloudflare viable
Are users geographically distributed?Cloudflare's anycast helpsRegional SFU may suffice
Is the application primarily audio?Cloudflare likely cheaperCost differences smaller
Do you need specific compliance certifications?Verify Cloudflare statusLess concern
Are you already on Cloudflare's platform?Integration is simplerEvaluate independently

Start with the simplest option that works. If your "real-time" feature involves text, JSON, or application state, Durable Objects with WebSockets almost certainly suffice. Don't add media infrastructure complexity for problems that don't require it.

Prototype before committing. The SFU’s shared 1,000 GB monthly allowance supports meaningful testing. Build a representative call, test across real networks, and measure latency, quality and billable traffic. Real-time media depends heavily on client behaviour and network conditions.

Plan for operational complexity. Real-time media is harder to operate than request/response applications. Budget for monitoring, debugging tools, and user support for "the call isn't working" issues. This complexity exists regardless of provider choice.

What comes next​

A call has more state than its media streams: membership, permissions, recordings and billing records each need a home. Chapter 12 begins the storage section by choosing that home from consistency, access and retention requirements.