Building Remota: A WebRTC-Based Remote Desktop Tool from Scratch
Remote desktop tools like TeamViewer, AnyDesk, and RustDesk are products that most developers use without thinking about what makes them work. I decided to build one from scratch for an on-demand client. The result is Remota, a remote access tool that lets a controller view and control a participant's Windows desktop directly from a browser, with no account required, no installation on the controller side, and a native Rust application handling the heavy lifting on the participant side.
This article documents the full technical journey: the architecture decisions, the protocols involved, the problems encountered, and the specific solutions that made it work.
Why Build This
Existing remote desktop tools are heavyweight, require accounts, often have licensing restrictions for commercial use, and do not give you control over the infrastructure. The goal was a lightweight tool where a controller could share a link, a participant could click it, and a live remote desktop session would be established. No accounts. No persistent agents. No vendor dependency.
The constraint that made this interesting is the browser. The controller should need nothing installed. That means the video stream and the control channel both had to work inside a standard browser tab using only web APIs. WebRTC is the only viable answer for that.
Architecture Overview
Remota has three components:
The backend is a Node.js Express server with WebSocket signaling. It manages rooms, validates tokens, relays WebRTC offer/answer/ICE messages between parties, and terminates sessions. It does not handle media traffic.
The web frontend is a React application built with Vite and Tailwind CSS. It serves two distinct roles depending on who is using it. The controller opens the app, creates a connection, shares the link, and views the remote desktop in a video element. The participant opens the link, accepts the connection, and either shares their screen directly from the browser or launches the native desktop application.
The desktop endpoint is a native Windows application written in Rust. It captures the screen using the Windows Graphics Capture API, encodes frames as VP8 using libvpx, sends the video stream over WebRTC, and injects mouse and keyboard input from the controller using Win32 SendInput.
The signaling server runs on an EC2 instance alongside a coturn TURN server for NAT traversal. The frontend is deployed on Vercel. The database is PostgreSQL on Supabase.
The Session Model
A session is represented as a Room in the database. Rooms have a token, a status, a creation timestamp, and an expiry timestamp. The token is a 43-character base64url string generated from 32 cryptographically random bytes. This is what goes in the shareable link.
The session lifecycle follows this sequence: the controller creates a room via a POST request, receives the token, and shares the link. The participant opens the link, validates the room, accepts the connection on a consent screen, and either starts browser share mode or launches the desktop app. The server tracks room status through WAITING, CONNECTING, ACTIVE, and ENDED states.
One of the early design mistakes was distributing termination authority across too many components. The original implementation had the controller navigating away on WebRTC disconnection, the participant cleaning up on WebSocket close, and the server terminating on timeout, all independently. This caused a persistent pattern of regressions: fixing one disconnect scenario would break another.
The correct model is that the server is the sole termination authority. Clients send a terminate request. The server broadcasts terminate to all parties and closes the room. Clients only navigate away after receiving that message from the server. WebRTC disconnected state is transient and should be ignored. Only failed state is terminal, and even then the client should request termination from the server rather than cleaning up unilaterally.
WebRTC Signaling
WebRTC requires a signaling channel to exchange SDP offers, answers, and ICE candidates before the peer-to-peer connection can be established. The signaling server uses the ws library in Node.js attached to the same HTTP server as Express.
Each connected client joins a room with a role: controller or participant. The server stores a map of rooms to connected clients. When the participant joins, the server notifies the controller with a participant_joined message. The controller creates an RTCPeerConnection, adds a recvonly video transceiver and a DataChannel, and sends an offer. The server relays it to the participant. The participant creates its own RTCPeerConnection, sets the remote description, adds its tracks, creates an answer, and sends it back. ICE candidates are relayed in both directions.
A keepalive mechanism was necessary because WebSocket connections over nginx would silently drop after periods of inactivity once the initial signaling exchange was complete. The server sends a native WebSocket ping every 30 seconds to all connected clients. Clients that do not respond with a pong are terminated. The browser-side signaling client also sends a JSON ping message every 25 seconds to prevent proxy timeouts, since browsers do not expose native WebSocket ping frames.
Screen Capture on Windows
The Windows Graphics Capture API, exposed through the windows-capture crate, captures the primary monitor using the WinRT ScreenCaptureKit equivalent. It delivers frames as BGRA pixel data via a callback on an internal capture thread. The callback receives a Frame object, and the pixel data is extracted using the as_nopadding_buffer method which strips row padding and returns a packed BGRA byte slice.
An important detail: as_nopadding_buffer takes a mutable Vec reference and returns a borrowed slice pointing into the frame buffer. It does not write into the Vec. An early version of the code was passing the Vec and ignoring the return value, then sending the empty Vec to the encoder. The encoder received zero bytes and panicked on the first array access. The fix was to use the returned slice directly and copy it into an owned Vec.
The capture API runs on a background OS thread. The frames are sent to the encoding thread via a tokio mpsc channel using the blocking_send method, since the capture callback is synchronous.
VP8 Encoding
The VP8 encoder is built on the vpx-encode crate which wraps libvpx through the env-libvpx-sys bindings. The encoder takes planar I420 data, so each BGRA frame must be converted before encoding.
The BGRA to I420 conversion applies BT.601 coefficients. For each pixel, the Y luma value is computed from all three color channels. U and V chroma values are computed for every 2x2 block of pixels by averaging the BGRA values of the four pixels in the block.
The libvpx encoder requires the timebase to be set at initialization. The initial implementation used a timebase of 1/30, representing one tick per frame at 30 frames per second. This caused a VPX_CODEC_INVALID_PARAM error on initialization with libvpx 1.14 built by vcpkg on the GitHub Actions runner. The correct timebase is 1/1000000, representing microseconds. The presentation timestamp passed to each encode call is then the frame index multiplied by 1000000 divided by the target frame rate.
The encoder is not Send because it contains raw C pointers. This means it cannot live inside a tokio task across await points. The solution is to run the encoder on a dedicated OS thread and communicate with the async WebRTC write task via a channel. The OS thread uses Handle::current().block_on() to drive receives from the tokio channel. This requires capturing the runtime handle before spawning the thread, since inside the thread there is no active runtime context.
The Keyframe Timing Problem
This was the most subtle bug in the entire project and took the longest to diagnose.
The encoding loop starts when the desktop application launches, before the WebRTC offer has been received. The first keyframe, a large packet typically around 140 kilobytes, is encoded and passed to TrackLocalStaticSample::write_sample within the first second of the application running.
However, webrtc-rs silently discards write_sample calls when no RTP sender is bound to the track. The RTP sender is only bound after set_local_description is called during the offer/answer exchange. By the time the WebRTC connection reaches the Connected state, the only keyframe that was produced has already been discarded.
The browser then receives only delta frames. A VP8 decoder cannot reconstruct video from delta frames without a preceding keyframe. The result is a completely black video element even though the pipeline is otherwise working correctly: the encoder is producing valid packets, write_sample is succeeding, and the browser reports receiving RTP packets with a non-zero packet count.
The fix is to trigger a forced keyframe immediately when the WebRTC connection reaches Connected state. This is implemented with an AtomicBool shared between the main async loop and the encoding thread. When the Connected state fires, the main loop sets force_keyframe to true. The encoding thread checks this flag before each encode call. When it is set, the encoder is reset by clearing the last known dimensions, which causes the encoder to be rebuilt on the next frame, producing a fresh keyframe that the browser can decode.
Browser Mode vs Desktop Mode
One design decision that added significant complexity was supporting two participation modes.
In browser mode, the participant uses getDisplayMedia to capture their screen directly in the browser and streams it over WebRTC. No installation is required. The controller can view the screen but cannot inject input because browsers do not expose OS-level input APIs. This mode works on desktop browsers but not on mobile, since Chrome on Android does not support getDisplayMedia.
In desktop mode, the native Rust application handles capture and input injection. The controller gets full mouse and keyboard control.
A third mode was added: viewer mode. A viewer can join a session to watch without sharing anything and without the ability to control. The signaling server was extended with an observer role for this purpose. Observer clients receive terminate messages from the server but their disconnection does not trigger the grace period logic that would terminate the room.
The browser mode introduced an audio problem that took several iterations to resolve. The participant's microphone is captured via getUserMedia and added to the stream alongside the video track. The controller can hear the participant. The controller can also speak using a mic captured via getUserMedia, and the audio is sent to the participant over the same peer connection. Echo cancellation is enabled on both sides using the echoCancellation constraint.
The issue is that the participant's mic starts unmuted by default if track.enabled is not explicitly set to false before adding the track to the stream. The fix is to set micTrack.enabled = false immediately after getUserMedia returns and before the track is added.
Touch Support on Mobile
The controller interface needed to work on Android Chrome, where the controller might be viewing a remote desktop from their phone.
Desktop browsers use mouse events. Mobile browsers use touch events. The controller video element handles pointer events for desktop and touch events for mobile. The touch handling required mapping gestures to remote control messages:
A single tap maps to a left click. A double tap within 300 milliseconds maps to a double click. A long press held for 600 milliseconds maps to a right click. A single finger drag that exceeds 8 pixels of movement maps to mouse down, a series of mouse moves, and mouse up. A two-finger drag maps to scroll by computing the midpoint delta between consecutive touch move events.
The keyboard input on mobile required a hidden input element that is focused when a keyboard button is pressed in the controller UI. The browser's software keyboard appears. Input events on the hidden element are forwarded as keyboard messages over the WebRTC DataChannel.
A significant mobile-specific problem was video autoplay. Mobile browsers block autoplay for video elements that have unmuted audio. The solution is to always start the video element muted, call play, and then set muted to false after play resolves. Browsers allow muted video autoplay in all cases. If the video has a separate audio track, that audio is placed on a hidden audio element rather than the same video element, since the audio element does not have the same autoplay restrictions once a user gesture has occurred.
The Session Lifecycle Regressions
The most painful part of building Remota was not any single technical problem. It was a pattern of regressions caused by an insufficiently designed session termination model.
The original implementation had termination decisions scattered across seventeen independent code paths. The controller navigated away on WebRTC disconnected state. The participant cleaned up on WebSocket close. The server terminated rooms immediately when the participant WebSocket closed. Various cleanup functions were called from multiple places with different assumptions about what had already been cleaned up.
Each fix to one scenario broke another. Adding a grace period on the server to handle network drops caused the controller UI to freeze for fifteen seconds when the participant genuinely left, because the controller had no handler for the participant_left message. Adding a participant_left handler caused reconnection to fail. Removing the beforeunload listener fixed spurious leave dialogs but meant refreshing the page appeared to do nothing until the grace period expired.
The root problem was that no single component was authoritative. The correct architecture, arrived at after a full audit of every termination path, is:
The server is the only entity that terminates a session. Clients request termination by sending a terminate message. The server broadcasts terminate to all parties and closes all connections. Clients only navigate away on receiving the server's terminate message. WebRTC disconnected state is transient and is ignored. WebRTC failed state causes the client to request termination from the server. WebSocket close triggers a ten-second grace period on the server. If the client reconnects within that window, the timer is cancelled. If not, the server terminates the room.
With this model, every disconnect scenario follows the same path through the server, and fixes to one scenario cannot break another.
Input Injection
Mouse and keyboard input is sent over the WebRTC DataChannel as JSON messages using a platform-independent protocol. Mouse move events include normalized coordinates in the range 0.0 to 1.0, which are scaled to the actual screen dimensions on the participant side. This allows the remote display to be rendered at different sizes on the controller without changing the coordinate protocol.
The desktop application dispatches these messages using the enigo crate, which wraps Win32 SendInput. All supported mouse operations work: movement, left click, right click, middle click, double click, scroll, and drag. Keyboard support covers standard characters, modifier keys, function keys, arrow keys, navigation keys, and special keys. The key mapping converts JavaScript KeyboardEvent.key strings to enigo Key enum values.
A critical safety requirement is that all held modifier keys are released on disconnect. If the connection drops while a modifier key is held down, the participant's computer would have stuck keys. The enigo controller calls release on Shift, Control, Alt, and Meta unconditionally on any exit path.
Protocol Handler Registration
To allow the participant web page to launch the desktop application directly via a link, the application registers the remota:// URI scheme on first run. This is done by writing keys to HKEY_CURRENT_USER\Software\Classes\remota in the Windows registry using the winreg crate. Using the current user hive rather than the local machine hive means no administrator privileges are required.
When the user clicks the Launch Remota Desktop button on the participant page, the browser opens a remota://session/TOKEN URI. Windows looks up the handler, finds the registered executable, and launches it with the URI as the first argument. The application parses the token from the URI and proceeds directly to connection without prompting the user.
The backend URL, TURN server credentials, and other configuration are baked into the binary at compile time using Cargo environment variables read in build.rs. The end user never needs to configure anything.
TURN Server and NAT Traversal
WebRTC attempts a direct peer-to-peer connection first using STUN to discover public IP addresses. When a direct connection cannot be established due to symmetric NAT or firewall restrictions, traffic is relayed through a TURN server.
Remota uses coturn deployed on the same EC2 instance as the Node.js backend. The STUN and TURN server runs on port 3478 over both TCP and UDP. The relay port range is 49152 to 65535. The WithoutBorder setting in the capture configuration prevents the yellow highlight border that Windows draws around any screen being captured.
The coturn configuration uses the long-term credential mechanism with a static username and password. These credentials are baked into both the frontend and the desktop binary at build time. The frontend reads them from Vite environment variables at build time. The desktop binary reads them from Cargo environment variables.
Deployment
The backend runs on an EC2 t4g.small instance, a Graviton2 ARM machine with 2 vCPUs and 2 gigabytes of RAM. Node.js, coturn, and nginx all run on this instance. PM2 manages the Node.js process in cluster mode. nginx handles TLS termination and proxies HTTP and WebSocket traffic to Node.js, with the proxy_read_timeout set to 3600 seconds to support long-lived WebSocket connections.
The PostgreSQL database is a Supabase instance. Prisma handles migrations and query building. The migration connection uses the direct URL to the PostgreSQL instance rather than the PgBouncer pooler, because PgBouncer in transaction mode does not support the advisory locks that Prisma needs for migrations.
The frontend is deployed on Vercel with a vercel.json rewrite rule that sends all requests to index.html for client-side routing. Vite environment variables containing the backend URL and TURN credentials are set in the Vercel project settings and baked into the static bundle at build time.
The desktop binary is built on GitHub Actions using a windows-latest runner. The build installs libvpx via vcpkg, copies vpx.lib to libvpx.lib because env-libvpx-sys expects the latter name on Windows, sets the required environment variables, and runs cargo build in release mode. The ffi-generate feature is used to regenerate the libvpx FFI bindings at build time from the installed headers, avoiding ABI version mismatches between the pre-generated binding files and the installed library version.
What Works
The full session flow works end to end. The controller creates a connection, shares a link, and the participant accepts. In browser mode, the participant shares their screen from the browser and the controller sees it in real time with audio. In desktop mode, the native application captures the screen, streams VP8 video, and the controller can move the mouse, click, scroll, type, and use keyboard shortcuts. Input injection works correctly and all modifier keys are released on disconnect. Sessions terminate cleanly from either side.
What Does Not Work Yet
The macOS desktop endpoint code is written but has not been tested because macOS builds require Xcode which only runs on macOS. The GitHub Actions macOS runner would handle this when the workflow is extended.
The desktop VP8 video stream has the correct pipeline but the browser overlay debug tooling was still in place at the time of writing. Once the diagnostic logging is removed and the binary is distributed normally the video should render as expected based on the confirmed RTP packet flow.
The Android and iOS participant applications are roadmap items. The platform restrictions are significant: Android requires the MediaProjection API and the AccessibilityService for input injection, neither of which is available from a browser. iOS has similar restrictions with ReplayKit. The Remota control protocol is designed to be platform-independent so future mobile clients can implement the same message format.
Lessons Learned
Several things became clear over the course of building this.
Session lifecycle design needs to happen before any code is written. Distributing termination decisions across multiple components creates a system where every fix introduces a new regression. The server-as-sole-authority model is not obvious but it is the only model that is consistent.
VP8 keyframe timing is not obvious. The fact that write_sample calls are silently discarded before the RTP sender is bound is not documented prominently. Any application that starts encoding before the WebRTC connection is established must force a keyframe on connection.
WebRTC disconnected state is not terminal. The specification defines it as a state the connection may recover from. Treating it as terminal causes sessions to end on brief network interruptions. Only failed state is truly unrecoverable from the browser's perspective.
libvpx version mismatches between compile-time bindings and runtime library cause silent initialization failures with error code 3. The ffi-generate feature solves this by regenerating bindings from the actual installed headers at build time.
The Rust async runtime does not allow non-Send types to cross await points. The VP8 encoder contains raw C pointers and is not Send. The solution is to confine the encoder to a dedicated OS thread and communicate with the async runtime via channels, capturing the runtime handle before spawning the thread.
Let me sign out here
Remota is a functional remote desktop tool built on standard protocols: WebRTC for media and data transport, WebSocket for signaling, and Win32 for native OS integration. The web controller requires no installation. The participant installs a single binary that registers a protocol handler and connects automatically from a browser link.
The hardest problems were not the ones that seemed hardest at the start. Screen capture and VP8 encoding were straightforward once the correct library versions and configurations were identified. The genuinely hard problems were session lifecycle correctness, keyframe timing, and the interaction between mobile autoplay policies and live media streams.
The source code is available at github.com/KingDavidJnr/remota.
Comments
Post a Comment