Writing
RustJul 12, 2026/12 min read

Two Ways to Handle a Thousand Clients: Multithreading vs. I/O Multiplexing

Two Ways to Handle a Thousand Clients: Multithreading vs. I/O Multiplexing

A hands-on comparison in Rust — every single line of code explained, no scary hand-waving, mild jokes included.


So there I was, feeling pretty good about myself. I had written a TCP server in Rust. It listened on a port. It accepted a connection. It echoed back whatever I typed. Beautiful. I was basically a systems programmer now.

Then I opened a second terminal, connected a second client, and... nothing. My server just sat there, loyally staring at client #1, completely ghosting client #2.

That's when I slammed into the question every network programmer eventually faces:

My server can talk to one client. How do I make it talk to thousands — at the same time?

It turns out there are two classic answers, and understanding the difference between them is one of those "ohhh, the whole field suddenly makes sense" moments. Nginx, Node.js, Tokio — once you see these two models, you understand why each of them is built the way it is. It's like learning that every magic trick is either a hidden compartment or a distraction.

So here's the plan: we'll build the same tiny program twice — a TCP echo server (it sends back whatever you type, like a very polite parrot). The logic stays identical both times. The only thing that changes is how it juggles many clients at once. And we'll walk through every line, because the details are where the understanding lives.

Grab a coffee. Let's go.


The problem, concretely

A basic TCP server does four things:

  1. Listen on a port.
  2. Accept a connection.
  3. Read from it, write back to it.
  4. Repeat.

The trap is step 3. read() blocks — if the client hasn't sent anything yet, your program freezes on that line and waits. Maybe forever. Your server isn't crashed, it isn't broken, it's just... standing there. Refreshing its inbox. So while you're stuck waiting on client A (who went to make tea), clients B, C, and D can't get a word in. One slow client freezes everyone.

There are two famous escapes from this trap:

  • Multithreading — give every client its own thread. If one thread blocks, the OS simply runs a different one.
  • I/O multiplexing — keep a single thread, make sure it never blocks on any one client, and instead ask the operating system: "Which of these clients is ready right now?"

Here's an analogy to keep in your head for the rest of the post. Picture a restaurant:

Multithreading hires one waiter per table. Each waiter is allowed to stand around doing absolutely nothing while their customer reads the menu for the fourth time — it doesn't matter, every table has its own personal waiter.

I/O multiplexing hires one very sharp waiter for the entire room, who never lingers at any single table. They scan the room and go only to tables that are actively waving.

More workers, or one smarter worker. That's the whole debate, and roughly forty years of server architecture. Let's build both.


Approach 1: Multithreading — one thread per connection

The idea: every time a client connects, spawn a brand-new OS thread just for them. Inside that thread we can write simple, blocking code — because if this thread freezes on read(), the OS scheduler just switches to another thread. Blocking becomes harmless.

Here's the entire server. Standard library only — zero dependencies.

use std::io::{Read, Write};
use std::net::{TcpListener, TcpStream};
use std::thread;
 
fn main() -> std::io::Result<()> {
    let listener = TcpListener::bind("127.0.0.1:8080")?;
 
    for stream in listener.incoming() {
        let stream = stream?;
        thread::spawn(move || handle(stream));
    }
    Ok(())
}
 
fn handle(mut stream: TcpStream) {
    let mut buf = [0u8; 1024];
    loop {
        match stream.read(&mut buf) {
            Ok(0) => return,
            Ok(n) => { let _ = stream.write_all(&buf[..n]); }
            Err(_) => return,
        }
    }
}

Twenty-odd lines. Now let's take it apart.

Line by line

We'll follow the code the way it actually runs: start in main(), where connections arrive, then zoom into handle(), where each client lives.

use std::io::{Read, Write};

This imports the Read and Write traits. In Rust, methods like .read() and .write_all() live on traits, and you can only call them if the trait is in scope. Delete this line and stream.read(...) won't even compile — a classic beginner stumble.

use std::net::{TcpListener, TcpStream};

Two actors: TcpListener sits on a port waiting for connections; TcpStream represents one established connection to one client. You read from and write to a TcpStream.

use std::thread;

The standard-library threading module — the star of this approach.

Part 1: main() — the front door

fn main() -> std::io::Result<()> {

The entry point returns io::Result<()> so we can use the ? operator: any I/O error bubbles up and exits the program with a message.

    let listener = TcpListener::bind("127.0.0.1:8080")?;

Claim port 8080 on localhost and start listening. If the bind fails (say, the port is already in use), ? returns that error immediately.

    for stream in listener.incoming() {

incoming() is an iterator that blocks until a client connects, then yields the new connection. This loop runs once per client, forever.

        let stream = stream?;

Each item is actually a Result<TcpStream>, because accepting a connection can fail. ? unwraps the success case.

        thread::spawn(move || handle(stream));

This one line is the entire trick. thread::spawn creates a fresh OS thread and runs handle(stream) inside it. The move keyword transfers ownership of stream into the closure, so the new thread fully owns its client.

And here's the key: main does not wait for handle to finish. It immediately loops back to incoming() to accept the next client. So with 500 clients connected, we have 500 threads, each blissfully blocking on its own read(), none of them disturbing the others.

So main is just a hiring manager: client walks in, spawn a worker, hand the client over, next please. But what does each worker actually do? Let's follow one client into its thread.

Part 2: handle() — one client's whole life

fn handle(mut stream: TcpStream) {

The function that serves exactly one client, taking full ownership of that client's stream. It's mut because reading and writing mutate the stream's internal state (buffers, cursor position).

    let mut buf = [0u8; 1024];

A fixed 1024-byte buffer on the stack. 0u8 means "the number zero, as an unsigned 8-bit byte," so this is 1024 zeroed bytes. Every read() dumps incoming data here.

    loop {

An infinite loop, on purpose. We want to keep serving this client — read a message, echo it, read the next — until they hang up. Without it, we'd handle one read and drop the connection.

        match stream.read(&mut buf) {

Try to read from the client. Two crucial facts about read():

  1. It blocks — if the client sent nothing, this line waits, possibly forever. In this design, that's fine: only this one thread is parked.
  2. It returns the number of bytes actually read, which drives the three arms below.
            Ok(0) => return,

Zero bytes is TCP's polite way of saying "the client closed the connection." We return, the function ends, the thread ends. Clean exit.

            Ok(n) => { let _ = stream.write_all(&buf[..n]); }

We received n bytes. &buf[..n] slices the buffer down to exactly the bytes that arrived (ignoring the untouched zeros after them), and write_all sends them straight back. That's the echo. The let _ = deliberately discards the result — a production server would handle a write error, but we're keeping the demo minimal.

            Err(_) => return,

Any read error — client crashed, connection reset — and we give up on this client, ending the thread. The _ says "I don't care which error."

Checkpoint: what this design costs

Threads aren't free. Each OS thread reserves its own stack — often a couple of megabytes — and the OS scheduler pays a tax switching between them. At a few hundred connections, nobody notices. At tens of thousands, you're burning gigabytes of RAM on stacks that are 99% empty, and your CPU spends more time deciding who to run than actually running anyone. It's the restaurant with 10,000 waiters: the kitchen is fine, but nobody can move.

This is the historically famous C10k problem: how do you serve 10,000 concurrent connections on one machine? Thread-per-connection buckles here — and that pain is exactly what motivated the second approach.

Multithreading in one line: simple to write, blocking is allowed, but every connection costs a whole thread.


Approach 2: I/O multiplexing — one thread, many clients

Now the clever one. Instead of one thread per client, a single thread manages every client at once. Yes, one. No, it's not a typo.

The magic ingredient is a family of system calls — epoll on Linux, kqueue on macOS, IOCP on Windows — that let you hand the OS a big list of sockets and say:

"Go to sleep. Wake me the moment any of these becomes ready. And tell me which ones."

Basically, we stop bugging every socket individually ("anything? anything? anything?") and instead let the OS tap us on the shoulder when there's actual news.

That's multiplexing: one thread, spread across many connections. Two rules make it work:

  1. Every socket is set to non-blocking — a read() with no data available returns a special error called WouldBlock instead of freezing.
  2. The thread only ever sleeps in one place: the "wake me when someone's ready" call.

Talking to raw epoll/kqueue/IOCP separately for each OS would be miserable, so we'll use the mio crate — a thin, cross-platform wrapper over all three. (Fun fact: mio is the foundation the entire Tokio async runtime is built on. You're about to touch the same machinery.)

Cargo.toml:

[dependencies]
mio = { version = "1", features = ["os-poll", "net"] }

The full program:

use std::collections::HashMap;
use std::io::{Read, Write};
use mio::net::{TcpListener, TcpStream};
use mio::{Events, Interest, Poll, Token};
 
const SERVER: Token = Token(0);
 
fn main() -> std::io::Result<()> {
    let mut poll = Poll::new()?;
    let mut events = Events::with_capacity(128);
 
    let mut listener = TcpListener::bind("127.0.0.1:8080".parse().unwrap())?;
    poll.registry().register(&mut listener, SERVER, Interest::READABLE)?;
 
    let mut connections: HashMap<Token, TcpStream> = HashMap::new();
    let mut next_token = 1;
 
    loop {
        poll.poll(&mut events, None)?;
 
        for event in events.iter() {
            match event.token() {
                SERVER => {
                    loop {
                        match listener.accept() {
                            Ok((mut stream, _addr)) => {
                                let token = Token(next_token);
                                next_token += 1;
                                poll.registry()
                                    .register(&mut stream, token, Interest::READABLE)?;
                                connections.insert(token, stream);
                            }
                            Err(ref e) if e.kind() == std::io::ErrorKind::WouldBlock => break,
                            Err(e) => return Err(e),
                        }
                    }
                }
                token => {
                    if let Some(stream) = connections.get_mut(&token) {
                        let mut buf = [0u8; 1024];
                        let mut closed = false;
                        loop {
                            match stream.read(&mut buf) {
                                Ok(0) => { closed = true; break; }
                                Ok(n) => { let _ = stream.write_all(&buf[..n]); }
                                Err(ref e) if e.kind() == std::io::ErrorKind::WouldBlock => break,
                                Err(_) => { closed = true; break; }
                            }
                        }
                        if closed {
                            let mut stream = connections.remove(&token).unwrap();
                            poll.registry().deregister(&mut stream)?;
                        }
                    }
                }
            }
        }
    }
}

Bigger, yes. But the payoff for taking it slowly is understanding how essentially every high-performance server on the planet actually works.

Line by line

use std::collections::HashMap;

All active connections will live in a HashMap. This is a philosophical shift: with one thread, we are now responsible for remembering every client. In Approach 1, each thread's stack did that bookkeeping for us, invisibly.

use std::io::{Read, Write};

Same traits as before, for .read() and .write_all().

use mio::net::{TcpListener, TcpStream};

Careful — these are mio's versions, not the standard library's. They look nearly identical but come pre-configured as non-blocking and ready to register with the event system.

use mio::{Events, Interest, Poll, Token};

The four core mio types, worth memorizing:

  • Poll — the event notifier. The thing we ask, "who's ready?"
  • Events — a container that gets filled with "these sockets are ready" notifications.
  • Interest — what we care about on a socket: readability, writability, or both.
  • Token — a small integer ID we pin to each socket, so when one wakes us up we know which one it was. Think of it as a coat-check ticket for connections.
const SERVER: Token = Token(0);

We reserve token 0 for the listener itself. When poll says "token 0 is ready," it means a new client wants to connect — as opposed to an existing client sending data.

    let mut poll = Poll::new()?;

Create the poll instance. Under the hood, this opens an epoll/kqueue/IOCP handle from the OS.

    let mut events = Events::with_capacity(128);

Pre-allocate room for up to 128 readiness notifications per wake-up. If more than 128 sockets are ready simultaneously, the extras just arrive on the next loop iteration — nothing is lost.

    let mut listener = TcpListener::bind("127.0.0.1:8080".parse().unwrap())?;

Bind port 8080, as before. .parse() converts the string into the SocketAddr that mio's bind expects. The unwrap() is fine here — the address literal is obviously valid.

    poll.registry().register(&mut listener, SERVER, Interest::READABLE)?;

Register the listener with the poll. In plain English: "OS, watch this listener. When it becomes READABLE — meaning a connection is waiting to be accepted — wake me up, and label the event with the SERVER token." This is how the listener joins the pool of things we're multiplexing over.

    let mut connections: HashMap<Token, TcpStream> = HashMap::new();

Our registry of live clients, keyed by token. When an event fires for token 7, we look up connections[&Token(7)] to find the actual socket.

    let mut next_token = 1;

A counter for handing out unique tokens. Starts at 1 because 0 belongs to the listener.

    loop {

The event loop — the heartbeat of the whole server. Every iteration: sleep until something is ready, then handle exactly the things that became ready.

        poll.poll(&mut events, None)?;

The single most important line in the program. It blocks the thread until at least one registered socket is ready, then fills events with the list of ready ones. None means "no timeout — sleep as long as it takes."

This is the only place the entire server is allowed to block. Parked here, the thread uses zero CPU — it's not busy-waiting, not polling in a hot loop, it is genuinely asleep, dreaming of packets. The instant any client sends data — or a new one connects — the OS wakes us with a precise list of exactly who needs attention. We never waste a single cycle checking idle sockets. This is our one sharp waiter, scanning the room.

        for event in events.iter() {

Walk through every socket that became ready this round. Could be one; could be fifty.

            match event.token() {

Which socket woke us? The token tells us. This is exactly why we labeled everything at registration time.

                SERVER => {

Token 0: the listener is ready — one or more new clients are waiting to be accepted.

                    loop {
                        match listener.accept() {

We loop on accept() because a single wake-up might mean several clients queued up at once. Keep accepting until the queue is empty.

                            Ok((mut stream, _addr)) => {

accept() succeeded, handing us a new client stream (and its address, ignored via _addr).

                                let token = Token(next_token);
                                next_token += 1;

Mint a fresh, unique token for this client and bump the counter for the next one.

                                poll.registry()
                                    .register(&mut stream, token, Interest::READABLE)?;

Add the newcomer to the multiplexing pool. From now on, poll also watches this socket and will wake us — with this token — whenever the client sends data. This is how the pool grows dynamically, one registration at a time.

                                connections.insert(token, stream);

Store the socket in our HashMap so we can find it when its token fires later. (mio wants the socket registered before it moves into the map — hence the ordering.)

                            Err(ref e) if e.kind() == std::io::ErrorKind::WouldBlock => break,

The non-blocking signature. When no more connections are waiting, accept() doesn't freeze — it returns WouldBlock, which really means "nothing left right now, come back later." That's our cue to break out of the accept loop. This is a normal, expected outcome, not a real error.

                            Err(e) => return Err(e),

Any other error is genuine trouble, so we propagate it up.

                token => {

The token wasn't 0, so it belongs to an existing client who just sent data. The variable token binds to whichever client it was.

                    if let Some(stream) = connections.get_mut(&token) {

Look that client up in the map. The if let Some(...) guards against the rare case where it's already been removed. get_mut because we're about to both read and write.

                        let mut buf = [0u8; 1024];
                        let mut closed = false;

A read buffer, plus a flag we'll raise if this client turns out to have disconnected.

                        loop {
                            match stream.read(&mut buf) {

Drain the socket. Since it's non-blocking, we keep reading until the OS says there's nothing more buffered. (One readiness notification can represent more bytes than fit in a single 1024-byte read.)

                                Ok(0) => { closed = true; break; }

Zero bytes — the client closed the connection. Flag it, stop reading.

                                Ok(n) => { let _ = stream.write_all(&buf[..n]); }

Got n bytes — echo them straight back, exactly like Approach 1.

                                Err(ref e) if e.kind() == std::io::ErrorKind::WouldBlock => break,

WouldBlock again — and here it's the good kind of stop: "you've read everything available for now." We break and move on to other clients. This is the essence of non-blocking I/O: instead of falling asleep when there's no data, we're told "come back later" and immediately go be useful somewhere else. WouldBlock sounds like an error, but in this world it's the most normal thing a socket can say to you. You'll learn to love it.

                                Err(_) => { closed = true; break; }

A real read error — treat the client as gone.

                        if closed {
                            let mut stream = connections.remove(&token).unwrap();
                            poll.registry().deregister(&mut stream)?;
                        }

If the client left, we clean up manually: remove it from the HashMap and deregister it so poll stops watching a dead socket. In Approach 1, the OS did this automatically when the thread ended. Here, cleanup is our job — the recurring trade-off of this model: more control, more responsibility.

Checkpoint: the mental model

One thread sits asleep in poll.poll(). The OS wakes it with a precise list: "these sockets need attention." The thread services exactly those, touches nothing idle, and goes back to sleep. There are no per-connection threads — a connection costs a HashMap entry and a socket. That is how a single thread juggles tens of thousands of clients.

But look at everything you now manage by hand: the connection map, the tokens, the WouldBlock dance, manual cleanup. All the complexity the OS scheduler quietly hid from you in Approach 1 is now sitting in your source code, in plain sight.

I/O multiplexing in one line: one thread scales to enormous numbers of connections — but you must never block, and you own all the bookkeeping.


Try it yourself

Both servers speak plain TCP, so you can poke them with netcat (or telnet):

# Terminal 1
cargo run
 
# Terminal 2
nc 127.0.0.1 8080
hello        # type this...
hello        # ...and the server echoes it back

Now the fun experiment: open several terminals with nc at once. Both servers handle them fine — but for very different reasons. In Approach 1, run ps -o thr -p <pid> (or check your process viewer) and watch the thread count climb with every connection. In Approach 2, it never moves. Same behavior, wildly different machinery.


Two honest caveats (so you don't ship a bug)

I kept these examples minimal for teaching. Two things a production server must handle that I deliberately glossed over:

1. Writes can "would-block" too. If a client is slow to receive, the OS send buffer fills up, and write_all in the multiplexed version can itself hit WouldBlock. A fully correct event loop buffers the unsent bytes and registers Interest::WRITABLE to finish the send later. I skipped that to keep the echo loop readable — but it is the number-one thing people get wrong when they graduate from toy servers.

2. These models aren't either/or. The best of both worlds is one event loop per CPU core: multiplexing within each thread, and multiple threads to light up every core. That's exactly what Nginx and Tokio do. Multiplexing solves "many connections per thread"; threads solve "use all my hardware." Combine them.


So which should you use?

MultithreadingI/O multiplexing
Threadsone per connectionone total (or one per core)
Blocking I/O✅ allowed❌ forbidden — must be non-blocking
Who tracks connectionsthe OS (thread stacks)you (explicit HashMap + tokens)
Cost per connectiona whole thread (~MBs of stack)a socket + a map entry
Scales to 10k+ connectionsstrugglesbuilt for it
Code complexitylowhigher (event loop + state machine)
Under the hoodOS schedulerepoll / kqueue / IOCP

A rough rule of thumb:

  • Reach for threads when connections are modest in number, or when each connection does heavy CPU work — threads let you actually use multiple cores for that work.
  • Reach for multiplexing when you have lots of mostly-idle connections doing small, fast operations — chat servers, proxies, anything where thousands of clients sit around quietly and occasionally say something short.

And honestly? Don't just take my word for it. Write both. Break both. Connect fifty nc sessions, kill one mid-message, watch what happens. The line-by-line understanding you build here is the exact foundation that makes Nginx, Node.js, and Tokio stop looking like magic and start looking like engineering.

Building the multiplexed version with your own hands is the moment the design of every high-performance server clicks into place. In the next post, I'll take everything we built here and show you how a very famous piece of software bet its entire architecture on one of these two models — and won big. Stay tuned. 👀

Thanks for reading. Now go bind a socket. 🦀