A walkthrough of akkovman/gchat-go: two goroutines per connection, ping-pong deadlines, and why http.Server.Shutdown leaves WebSockets running.
· 8 min read

How gchat-go Structures a WebSocket Hub in Go


Call http.Server.Shutdown on a server whose only long-lived connections are 500 WebSockets and it can return while all 500 are still open. The method closes listeners, closes idle HTTP connections, and waits for active HTTP connections, but it deliberately does not close or wait for hijacked connections such as WebSockets.

gchat-go makes that missing lifecycle explicit, which is most of why it’s worth reading. It’s a global anonymous chat server: pick a nickname, connect, and you’re in one room with everybody else. Every accepted client gets two goroutines and a buffered channel, all feeding a shared hub that must also participate in shutdown.

The design closely follows the Gorilla WebSocket chat example and puts it behind Gin. This article examines commit 6ca08c4, the repository state when this post was raised. It is small enough to trace end to end, including a few shutdown and locking details you would want to change before copying it into production.

Two goroutines per connection, one buffered channel

The core type is in internal/ws/client.go. A Client holds a pointer back to the hub, the *websocket.Conn, a nickname, and a sendBuffer chan outboundMessage created with capacity 256 in internal/ws/handler.go.

This design uses one readPump and one writePump because Gorilla supports one concurrent reader and one concurrent writer per connection. The application must ensure that no two goroutines call its write methods concurrently, and the same rule applies to its read methods. Here, normal outbound messages go through sendBuffer rather than calling WriteMessage from several goroutines.

The ownership pattern is useful beyond WebSockets: when a resource is not safe for concurrent use, route operations through one goroutine. A channel send is synchronized before its matching receive under the Go memory model. That safely publishes the value, but it does not make later unsynchronized mutations to shared objects inside that value safe.

Buffer size is a real policy decision. In this code, broadcasts use a non-blocking send. A client can accumulate up to 256 queued messages; when the buffer is full, the next broadcast closes its channel and removes it from the hub instead of holding up every other client.

There is a locking bug in that slow-client branch. broadcast holds h.mu.RLock() while deleting from h.clients, even though addClient and deleteClient mutate the same map under h.mu.Lock(). The branch must take the write lock (or move removals into a single hub-owner goroutine); otherwise concurrent hub activity can race with the map deletion.

Read deadlines and the pong handler

readPump sets three things before its loop:

c.conn.SetReadLimit(int64(env.GetEnvAsInt("MAX_MESSAGE_SIZE", 4096)))
c.conn.SetReadDeadline(time.Now().Add(pongWait))
c.conn.SetPongHandler(func(string) error {
    c.conn.SetReadDeadline(time.Now().Add(pongWait))
    return nil
})

All three are liveness machinery, and all three are easy to botch.

SetReadLimit caps a message at 4096 bytes by default. Gorilla sends a close message and returns ErrReadLimit when a peer exceeds it. That puts a useful bound on work and memory, but it belongs alongside origin checks, connection limits, and transport-level limits rather than standing in for them.

The read deadline is absolute. It is not a rolling idle timeout. ReadMessage fails the moment the clock passes it, chatty connection or not. The pong handler is the only thing pushing it forward: peer answers a ping, deadline jumps another pongWait into the future. Forget the handler and every connection dies after 50 seconds no matter how much traffic it’s carrying.

The other half sits in writePump, which runs a time.Ticker at pingPeriod, defined in internal/ws/client.go as (pongWait * 9) / 10. Pinging at 90% of the tolerance leaves the round trip time to land before the deadline expires. Set pingPeriod above pongWait and you’ve built a timer that disconnects perfectly healthy clients.

The asymmetry in deadlines is worth noticing. The read deadline moves on pongs, while the write path calls SetWriteDeadline(time.Now().Add(writeWait)) only before a ping. Text and close messages do not set a fresh deadline first, so writeWait is not a per-message guarantee. If every write must be bounded, set the deadline immediately before every write.

The read loop rejects malformed input by continuing, not closing

Inside the loop, readPump unmarshals into an anonymous struct with Type and Text fields and applies one guard:

if json.Unmarshal(raw, &in) != nil || in.Type != "message" || in.Text == "" {
    continue
}

Bad JSON, wrong type, or empty text: one branch, frame dropped, client stays connected. That is a defensible chat-server policy, although a production service may also want metrics or a threshold so a client cannot send malformed frames forever without any feedback.

The anonymous struct is a small thing I like. It’s declared inside the loop, used only there, never promoted to a package-level type nobody else needs. Go lets you describe a shape exactly where it matters, and most codebases forget that.

Cleanup lives in deferred functions

Both pumps open with a defer that handles teardown. readPump calls c.hub.WG.Done(), removes the client with c.hub.deleteClient(c), and closes the connection. writePump stops its ticker, calls WG.Done(), broadcasts a system “left” message, and closes the connection.

Both pumps close the same *websocket.Conn. Closing it from either side makes the other pump’s next network operation fail, and a second close is harmless here. There is one ordering detail to fix, though: each deferred function calls WG.Done() before finishing its remaining cleanup. That means WG.Wait() can return while a pump is still deleting the client, broadcasting the leave message, or closing the connection. Call Done after cleanup if the wait is meant to prove the goroutine has fully finished.

For the finer points of how defer interacts with named returns and evaluation order, we covered that in How to defer in Golang like a pro.

Graceful shutdown: http.Server.Shutdown doesn’t close WebSockets

The overall sequence in cmd/server/main.go is the right outline, but not the whole production-ready implementation.

Shutdown waits for active requests to finish. A hijacked connection, which is what a WebSocket is the instant the upgrade completes, is not an active request as far as the server is concerned. The net/http docs state it plainly: Shutdown does not attempt to close or wait for hijacked connections. It has no idea they exist.

gchat-go does three things once the signal lands:

ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if err := srv.Shutdown(ctx); err != nil {
    log.Println("Server Shutdown:", err)
}

hub.CloseAll()
log.Println("Closing active Websocket connections...")
hub.WG.Wait()

Shutdown stops new HTTP work and waits for ordinary active requests. hub.CloseAll() enqueues a WebSocket close message for each client, and hub.WG.Wait() waits for both pump counters. Because Done currently runs before cleanup completes, the wait does not quite prove that every connection has closed. It also has no timeout of its own: the five-second context applies only to srv.Shutdown, so a stuck write pump can still hold shutdown open indefinitely.

In internal/ws/handler.go, the handler calls h.WG.Add(2) before starting either goroutine:

h.WG.Add(2)
go client.writePump()
go client.readPump()

Add(2) sits before the go statements. Moving it inside the goroutines would let Wait observe zero before the scheduler runs them. The two Done calls balance the count, but they should be the final teardown step if Wait is supposed to cover all cleanup.

Signal handling is the textbook version: a buffered channel of size 1 handed to signal.Notify for SIGINT and SIGTERM. The buffer is load-bearing. signal.Notify never blocks on send, so an unbuffered channel with no receiver parked on it at that exact moment silently drops the signal.

Ban list: RWMutex around a map

internal/ip/ban.go is about twenty lines. A BanList wraps map[string]bool with a sync.RWMutex. Block and Unblock take the write lock, while Check takes the read lock.

The access pattern is read-heavy: Check runs on every connection attempt, while Block and Unblock are operator commands. An RWMutex permits concurrent checks, although the map lookup is so small that you should measure contention before assuming it beats a plain Mutex.

Those operator commands arrive from somewhere unusual. main.go runs a bufio.Scanner over os.Stdin, splits lines with strings.Fields, and dispatches on block and unblock. It is a compact manual control surface, but automation and authentication are easier to design around an admin API or a proper Cobra CLI. One small bug is worth noting: a blank input line makes the goroutine return, permanently disabling the console; it should continue instead.

Rejecting a client after the upgrade, not before

ServeWS in internal/ws/handler.go has a sequencing detail that slides right past you. The ban check runs before the upgrade and returns http.StatusForbidden. Nickname validation runs after upgrader.Upgrade succeeds, and rejections go out as WebSocket close frames with custom codes: InvalidNickname, NicknameIsTaken, ServerIsFull.

The split follows the protocol. Once the handshake completes, there is no HTTP response status left to send. A close frame is the structured error channel, and RFC 6455 reserves codes 4000-4999 for private use. A browser can read event.code in onclose and render the corresponding message.

The banned-IP case is checked first, while the handler can still return 403 Forbidden. It relies on ctx.ClientIP(), which in turn relies on router.SetTrustedProxies being configured correctly; this server defaults to the 172.18.0.0/16 Docker bridge range. Trust too broad a proxy range and a request arriving through a trusted address may be able to influence the forwarding headers used for the ban check.

What to take from it

The general lifecycle transfers to other long-lived connections: stop accepting work, notify each connection, bound its final writes, and wait for ownership goroutines to finish. WebSockets add their own ping/pong rules; server-sent events and gRPC streams need protocol-specific cancellation instead. The timeout values are workload decisions, not constants to copy.

Where this design starts to strain is the single room. One hub owns one broadcast fan-out and every client’s sendBuffer. Add rooms and the hub is the natural unit to partition, while shutdown needs a clear owner for every room and connection. gchat-go is worth reading because the shape is still small enough to see all at once—including the places where a clean pattern still needs careful locking and teardown.