RE:NODE

App hosting11 min read

Hosting SignalR apps: transports, proxies and scale-out

Run ASP.NET Core SignalR in production: transports and negotiation, WebSockets through a proxy, timeouts, auth tokens, and a Valkey backplane for scale-out.

0 readers

SignalR is the part of ASP.NET Core that keeps a connection open between server and browser so the server can push messages without being asked. Hosting it is easy on one server and gets interesting the moment anything sits in the middle. One instance behind a proxy needs three things: the proxy must pass the WebSocket upgrade, its idle timeout must be longer than SignalR's 15-second keep-alive, and authentication has to work without custom headers. Two or more instances need two more: sticky sessions, or WebSockets only with negotiation skipped, and a backplane such as Valkey so a message sent on one server reaches clients on the other. This post covers each of those, with the defaults that decide when connections drop.

What SignalR actually does on the wire#

A SignalR connection starts as an ordinary HTTP request. The client posts to /<hub>/negotiate, and the server replies with a connection token and the transports it supports. The client then picks the best one both sides accept:

TransportHow it worksNotes
WebSocketsOne full-duplex TCP connection after an HTTP upgradePreferred. Lowest overhead and latency
Server-Sent EventsA long-lived HTTP response for server-to-client, separate POSTs for client-to-serverFallback when WebSockets are blocked
Long PollingRepeated requests held open until there is dataLast resort. Works almost anywhere, costs the most

On top of the transport runs the hub protocol, JSON by default or MessagePack if you add Microsoft.AspNetCore.SignalR.Protocols.MessagePack on the server and the matching client package. MessagePack messages are smaller and cheaper to parse; JSON is easier to read in the browser's network tab.

The minimal server is two lines plus a hub class:

csharp
builder.Services.AddSignalR();var app = builder.Build();app.MapHub<ChatHub>("/hubs/chat");public class ChatHub : Hub{    public Task Send(string room, string text) =>        Clients.Group(room).SendAsync("message", Context.UserIdentifier, text);    public Task Join(string room) =>        Groups.AddToGroupAsync(Context.ConnectionId, room);}

And the browser side, using the @microsoft/signalr package:

javascript
import * as signalR from "@microsoft/signalr";const connection = new signalR.HubConnectionBuilder()  .withUrl("/hubs/chat")  .withAutomaticReconnect()  .build();connection.on("message", (user, text) => render(user, text));await connection.start();await connection.invoke("Join", "general");

withAutomaticReconnect() without arguments retries after 0, 2, 10 and 30 seconds, then gives up and fires onclose. Groups are not restored after a reconnect, because a reconnected client has a new connection ID; rejoin them in the onreconnected handler. That one detail explains most reports of "messages stop arriving after the wifi blipped".

Keep-alives and timeouts#

A SignalR connection that carries no messages is kept alive by pings. The defaults on both ends are tuned to each other:

SettingSideDefaultWhat it means
KeepAliveIntervalServer15 secondsServer pings an idle client this often
ClientTimeoutIntervalServer30 secondsServer drops a client it has not heard from in this long
HandshakeTimeoutServer15 secondsTime allowed for the initial handshake
keepAliveIntervalInMillisecondsJS client15,000Client pings the server this often
serverTimeoutInMillisecondsJS client30,000Client gives up if the server is silent this long

The rule is that each side's timeout should be about double the other side's keep-alive interval. If you raise KeepAliveInterval on the server, raise serverTimeoutInMilliseconds on the client to match, or clients will disconnect themselves for no reason.

csharp
builder.Services.AddSignalR(options =>{    options.KeepAliveInterval = TimeSpan.FromSeconds(15);    options.ClientTimeoutInterval = TimeSpan.FromSeconds(30);    options.MaximumReceiveMessageSize = 64 * 1024;});

MaximumReceiveMessageSize defaults to 32 KB per incoming message. Clients that send larger payloads are disconnected with an error that is easy to misread as a network problem. Raise it if you must, but large uploads belong on a normal HTTP endpoint, not in a hub method. MaximumParallelInvocationsPerClient defaults to 1, meaning a client's hub calls are processed one at a time; that is a sensible default because it stops one client monopolising the server, and a common surprise when a long-running hub method makes the next call from the same client wait.

The proxy has a timeout of its own. Nginx's proxy_read_timeout defaults to 60 seconds, which is comfortably longer than the 15-second keep-alive, so a default SignalR app survives a default nginx. Problems appear when someone raises the keep-alive interval to "reduce traffic" past the proxy's idle timeout, or when a load balancer somewhere in the path closes idle connections after 30 seconds. Keep the keep-alive well under the shortest idle timeout on the path.

WebSockets through a reverse proxy#

A WebSocket starts as an HTTP request with Upgrade: websocket and Connection: Upgrade headers. A proxy that does not pass those through, or speaks HTTP/1.0 to the backend, turns every connection into a failed upgrade. SignalR then quietly falls back to Server-Sent Events or long polling, the app keeps working, and nobody notices that every client is now holding extra requests open and the server is doing several times the work. Check the browser's network tab: the negotiate request is followed by a request to the hub URL with status 101 Switching Protocols when WebSockets are working.

For nginx the four lines are:

nginx
location /hubs/ {    proxy_pass http://127.0.0.1:5000;    proxy_http_version 1.1;    proxy_set_header Upgrade $http_upgrade;    proxy_set_header Connection "upgrade";    proxy_set_header Host $host;    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;    proxy_set_header X-Forwarded-Proto $scheme;    proxy_read_timeout 120s;}

WebSockets behind a reverse proxy covers the map block that handles both upgrade and ordinary requests in one location, and the errors you see when it is wrong. On the application side, enable forwarded headers so the app knows the original scheme and client address; ASP.NET Core behind a reverse proxy has the exact ForwardedHeadersOptions.

On RE:NODE, C# / .NET plans include a proxy slot: point an A record at the address shown and the certificate is issued and renewed automatically, with the client address in X-Forwarded-For. After deploying, open the network tab and confirm the 101 before you trust that WebSockets are what your users get. If you would rather connect straight to the app, the plan's allocated port is reachable directly and more ports can be added on the Network tab - in that case Kestrel terminates the connection itself.

Cross-origin clients need CORS with credentials, because SignalR sends cookies on negotiate:

csharp
builder.Services.AddCors(o => o.AddPolicy("app", p => p    .WithOrigins("https://app.example.com")    .AllowAnyHeader()    .AllowAnyMethod()    .AllowCredentials()));

AllowCredentials() cannot be combined with AllowAnyOrigin(); list the origins.

Authentication on a persistent connection#

Cookie authentication works with SignalR without any special handling: the browser sends the cookie on the negotiate request and on the WebSocket upgrade. Bearer tokens are harder, because the browser WebSocket API cannot set an Authorization header. The JavaScript client works around this by putting the token in the query string as access_token for WebSockets and Server-Sent Events, and the server has to be told to read it there:

csharp
builder.Services.AddAuthentication(JwtBearerDefaults.AuthenticationScheme)    .AddJwtBearer(options =>    {        // Authority, Audience or TokenValidationParameters as usual ...        options.Events = new JwtBearerEvents        {            OnMessageReceived = context =>            {                var token = context.Request.Query["access_token"];                if (!string.IsNullOrEmpty(token) &&                    context.HttpContext.Request.Path.StartsWithSegments("/hubs"))                {                    context.Token = token;                }                return Task.CompletedTask;            }        };    });

On the client, .withUrl("/hubs/chat", { accessTokenFactory: () => getToken() }). The factory is called on every connect and reconnect, so return a fresh token rather than a captured string.

Two consequences follow. Query strings appear in access logs, so make sure neither your proxy nor your request logging records the full URL for hub paths. And a token is checked when the connection is established, not on every message; a connection authenticated with a token that expires an hour later stays authenticated until it disconnects. If you need revocation to bite, close connections from the server when a user signs out, or set a maximum connection lifetime by tracking connections and aborting them with Context.Abort().

Clients.User(userId) sends to every connection a user has open - several tabs, a phone and a laptop. The user ID comes from IUserIdProvider, which by default reads the ClaimTypes.NameIdentifier claim. If your tokens put the ID in sub and claim mapping is off, Context.UserIdentifier is null and user-targeted messages go nowhere; register a custom provider that reads the claim you actually issue.

Sending from outside a hub#

Most real messages do not originate in a hub method. An order is paid in a webhook handler, a background job finishes, a price changes. Inject IHubContext<THub> anywhere in the app:

csharp
public class OrderNotifier(IHubContext<ChatHub> hub){    public Task OrderPaid(string userId, int orderId) =>        hub.Clients.User(userId).SendAsync("orderPaid", orderId);}

Hub instances are transient: one is created per method call and disposed afterwards, so state stored in a hub field is lost immediately. Store per-connection state keyed by Context.ConnectionId in a singleton service, or in Context.Items, which lives as long as the connection. Long-running work started from a hub should be handed to a background service rather than awaited inside the hub method, because of the one-invocation-at-a-time default above. .NET worker services and background jobs covers the queue pattern.

Memory and sizing per connection#

Each open connection holds buffers, a connection context, and whatever your application attaches to it - group memberships, user state, cached data. The baseline per idle WebSocket connection is small, but it grows with message sizes, with the number of groups, and above all with application state you keep per connection. The only reliable number is your own: open a known number of test connections with a script using the .NET client (Microsoft.AspNetCore.SignalR.Client) or a load-testing tool that speaks WebSockets, and read the memory graph before and after.

Practical guidance for a small server:

  • Watch the fallback transports. A thousand long-polling clients cost far more than a thousand WebSocket clients, because each poll is a full request cycle. If WebSockets fail through your proxy, fixing that is worth more than any hardware.
  • Cap upgraded connections. Kestrel's MaxConcurrentUpgradedConnections is unlimited by default. A cap turns an unexpected flood into refused connections rather than an out-of-memory stop. On RE:NODE a server that reaches its memory limit is stopped and restarted clean, which drops every connection at once and triggers a reconnect storm; leave headroom.
  • Stagger reconnects. After a restart, every client reconnects within seconds. The default retry delays already spread them a little; a custom IRetryPolicy with jitter spreads them more.
  • Blazor Server is SignalR. Every Blazor Server user is a SignalR connection with a circuit holding UI state in server memory, which is heavier than a typical hub connection. Blazor Server vs WebAssembly hosting goes through that sizing.

Scale-out with a Valkey backplane#

With two or more app instances, a client connected to instance A will not receive a message sent through IHubContext on instance B, because each instance only knows its own connections. A backplane fixes this by publishing every message to a shared pub/sub channel that all instances subscribe to.

bash
$ dotnet add package Microsoft.AspNetCore.SignalR.StackExchangeRedis
csharp
builder.Services.AddSignalR()    .AddStackExchangeRedis(builder.Configuration.GetConnectionString("Valkey")!, options =>    {        options.Configuration.ChannelPrefix = RedisChannel.Literal("myapp");    });

The connection string is a StackExchange.Redis configuration string, such as valkey.example.net:6380,password=...,abortConnect=false. Valkey speaks the Redis protocol, so the Redis backplane package works against it unchanged. Set ChannelPrefix when more than one app shares the same Valkey server, or their messages mix.

IHubContextpublishsubscribeWebhook handleron instance AApp instance ASignalR hubValkeypub/sub backplaneApp instance BSignalR hubBrowsersWebSockets
A message sent on one instance reaches clients on every instance

Know what a backplane does not do:

  • It does not remove the need for sticky sessions. The negotiate request and the subsequent connection must reach the same instance. Either configure the load balancer for session affinity, or have every client use WebSockets only with skipNegotiation: true and transport: signalR.HttpTransportType.WebSockets, which removes the separate negotiate step. The second option drops the fallback transports for clients behind networks that block WebSockets.
  • It does not store messages. If Valkey is unreachable, messages published during the outage are lost, and SignalR does not replay them. Messages that must arrive belong in a database or a queue, with SignalR as the notification that something new exists.
  • Every message goes to every instance. The backplane's throughput is the ceiling on the whole cluster's message rate. For most applications that ceiling is far away; for a very high fan-out workload it is the thing to measure.

Valkey's pub/sub model and how it differs from streams is covered in Valkey pub/sub and streams, and a single small Valkey server is plenty for a backplane - it holds no data, only messages in flight.

Troubleshooting#

`Error: Failed to start the transport 'WebSockets'` then a fallback. The upgrade is not reaching the app. Check the proxy's Upgrade and Connection headers and HTTP/1.1 to the backend.

`No Connection with that ID` after scaling to two instances. Negotiate and connect landed on different instances. Enable sticky sessions or skip negotiation with WebSockets only.

Connections drop every 60 or 100 seconds exactly. An idle timeout somewhere in the path is shorter than the gap between messages, and keep-alives are disabled or too infrequent. Restore the default 15-second keep-alive.

`Server timeout elapsed without receiving a message from the server.` The client's serverTimeoutInMilliseconds is shorter than twice the server's keep-alive, or the server is too busy to send pings - check CPU.

401 on the WebSocket but negotiate works. The token is sent in the query string for the socket and the server is not reading it. Add the OnMessageReceived handler.

Messages arrive on one tab but not another after a network change. The reconnected connection lost its groups. Rejoin in onreconnected.

FAQ#

Do I need a backplane for a single server?

No. One instance knows every connection, so Clients.All and Clients.Group already reach everybody. A backplane is only for two or more instances, and adds a dependency you would otherwise not have.

Can I use Valkey instead of Redis for the SignalR backplane?

Yes. The backplane package talks the Redis protocol through StackExchange.Redis, and Valkey implements that protocol. Point the connection string at the Valkey server and set a channel prefix if anything else uses it.

How many SignalR connections can one small server hold?

Thousands of idle WebSocket connections fit in a modest amount of memory; what limits you is the state your app keeps per connection and the message rate. Measure with a test client against your own hub rather than trusting a published figure.

Should I use Azure SignalR Service instead?

It takes connection handling off your servers and is the easiest way to very large connection counts, but it is an external dependency with its own pricing and your messages leave your infrastructure. For most apps on a few instances, a self-hosted backplane is enough.

Why does SignalR use long polling in production when it used WebSockets locally?

Something between the browser and the app is not passing the WebSocket upgrade - usually the reverse proxy, sometimes a corporate network. The app keeps working, which is why it goes unnoticed. Check for a 101 response in the network tab.


Comments

Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.

0/2000