RTFM · Tools
Shell pipelines: how Unix thinks
This chapter is not six funny symbols. It is the mental model underneath them: small programs that move bytes, connected by the shell, until bytes become a useful answer. Once that model is real to you, grep, sed, awk, sort, tee, find, xargs, logs, and half of Linux administration, stop being mysterious.
Every process has three streams
When a program runs, the kernel hands it three open file descriptors before main() executes a line. They are not magic: a file descriptor is just a small integer indexing the process's open-file table, and almost anything can sit behind it; a terminal, a regular file, a pipe, a socket, or /dev/null.
process
+-------------+
keyboard/file ---->| 0 stdin |
| |----> 1 stdout
| |
| |----> 2 stderr
+-------------+
| Descriptor | Name | Default connection | Intended cargo |
|---|---|---|---|
| 0 | stdin | keyboard (terminal) | data to consume |
| 1 | stdout | terminal | the work product: results, logs-as-data |
| 2 | stderr | terminal | diagnostics: errors, warnings, progress |
The separation of 1 and 2 is the whole trick. A well-behaved Unix tool puts RESULTS on stdout and NOISE on stderr, so results can be piped onward without diagnostics contaminating them. Commands that blur this deserve your suspicion.
Redirection is re-plumbing descriptors
Every redirection below is the shell rewiring one of those three integers before the command starts. The command itself cannot tell a terminal from a file, and that is the point.
command >output.txt # fd1 -> file (truncate)
command >>output.txt # fd1 -> file (append)
command 2>errors.txt # fd2 -> file
command >output.txt 2>errors.txt # fd1 and fd2 -> separate files
command >everything.txt 2>&1 # fd1 -> file, THEN fd2 -> wherever fd1 points
command 2>&1 >output.txt # fd2 -> terminal(!), then fd1 -> file
command </path/to/input # fd0 -> file
command >/dev/null # discard results
command 2>/dev/null # discard diagnostics
The two 2>&1 forms are where everybody trips, so compare them side by side. 2>&1 means 'duplicate descriptor 2 onto descriptor 1; wherever fd1 points RIGHT NOW'. It is a copy of a pointer, taken at that moment, not a permanent linkage.
command >everything.txt 2>&1
# step 1: fd1 --> everything.txt
# step 2: fd2 --> everything.txt (fd1 already points there)
#
# result: both streams land in the file. Correct.
command 2>&1 >output.txt
# step 1: fd2 --> (wherever fd1 points NOW: the terminal)
# step 2: fd1 --> output.txt
#
# result: errors stay on screen, results go to the file.
# The streams were crossed in the wrong order.
Bash offers >& as in command >&file and &>file for the common both-to-file case, but the numbered form above is portable and teaches the mechanism. Learn the mechanism once and the shorthand reads itself.
Pipes: stdout into stdin, nothing in between
stdout stdin
command A ------- pipe -------> command B
The | operator connects fd1 of the left command to fd0 of the right command; a kernel buffer in the middle, not a file. Two important truths follow. First, the shell creates the pipe, forks both processes, and hands each its end; the data never passes through any temporary file, which is why streaming a 40 GB log works in 4 GB of RAM. Second, stderr is NOT piped: command A's diagnostics still write to your terminal, by design, so errors stay visible while data flows.
Piping stderr too: three ways, one habit
command | consumer # data only; errors stay on terminal
command 2>&1 | consumer # data + errors, explicitly re-plumbed
command |& consumer # Bash shorthand for the line above
Use plain | when you want errors visible and data flowing. Use |& (or the explicit 2>&1 |) when the consumer should also see the diagnostics; for example when feeding a filter that selects ERROR lines from a noisy build.
|& is a Bash extension, not POSIX sh. It is correct on Saphira (bash ships), but scripts claiming portability should write 2>&1 | explicitly.
First pipelines, stage by stage
Count three words:
printf '%s\n' alpha beta gamma | wc -l
# -> 3
printf emits three lines (the %s\n format is applied per argument), wc -l counts newlines. Now a two-stage pipeline where every intermediate result is shown - this is the habit that makes pipelines legible:
printf '%s\n' banana apple banana pear | sort
# stage 1 output:
# apple
# banana
# banana
# pear
printf '%s\n' banana apple banana pear | sort | uniq -c
# stage 2 output:
# 1 apple
# 2 banana
# 1 pear
Read it again with the model in mind: bytes left to right, each program transforming a stream it knows nothing about. Nobody parsed anything; uniq merely noticed that equal lines are now ADJACENT, which is why sort | uniq is a pair, not a coincidence. uniq collapses adjacent duplicates only; fed unsorted input it reports duplicates that are not there.
tee: watch a pipeline without dismantling it
command1 |
+ tee /tmp/stage1 |
+ command2 |
+ tee /tmp/stage2 |
+ command3
tee copies its stdin to a file AND to its stdout, so you can inspect exactly what flowed between stages after the fact, while the pipeline runs untouched. When a five-stage pipeline misbehaves, tee turns 'something in here is wrong' into 'stage 2 receives garbage', which is a different, tractable problem.
Quoting and expansion: the shell edits your words first
Before any command runs, the shell performs expansion on its arguments: parameter expansion, globbing, command substitution, word splitting. The command never sees your quotes; they are instructions to the shell, consumed during parsing.
"$variable" # expands, then stays ONE word even if it contains spaces
'$literal' # nothing expands; dollars and stars stay literal
* # expands to matching filenames in the current directory
? # expands to any single character
$(command) # runs command, substitutes its stdout
Why quoting matters, demonstrated by the difference between two lines that look like twins:
file="my notes.txt; rm -rf ~/docs"
rm $file # shell splits on spaces -> rm my notes.txt; rm -rf ~/docs (disaster)
rm "$file" # one word, exactly the filename (correct)
Also keep two different kinds of 'input' separate in your head: TEXT FLOWING THROUGH A PIPE (streaming bytes between already-running processes) and COMMAND-LINE ARGUMENTS (a fixed list fixed before the program starts, subject to all the expansion rules above). sort <file reads a stream; sort file1 file2 takes arguments. Pipes carry streams; arguments are built once.
Grouping: { } versus ( )
{ command1; command2; } >combined.log
(
cd /somewhere
command
)
Braces { ...; } run the commands in the CURRENT shell; variables set inside survive, and the redirection applies to the group. Parentheses ( ... ) create a subshell: a child copy of the shell whose cd, variable changes and umask vanish when it exits. Use ( ) precisely when you want throwaway environment; the cd above cannot leak out and disturb the caller.
Command substitution
kernel="$(uname -r)"
printf 'Running kernel: %s\n' "$kernel"
$(...) is preferred over the old backtick syntax: it nests cleanly, reads cleanly, and its failure modes are quieter. When the substitution produces multiple words, the shell splits them; quote "$(...)" when you want one word, exactly as with variables.
Here-documents and here-strings
cat <<'EOF'
literal text
$nothing expands here
EOF
cat <<EOF
the kernel is $(uname -r) # expansion happens
EOF
Quoting the delimiter ('EOF') turns expansion OFF for the whole block; leaving it bare turns expansion ON. This distinction matters hugely the day you use a heredoc to generate a config file: one form writes literal $ variables for the target program, the other bakes in the generating shell's values. Bash also offers here-strings (verified in the local manual as [n]<<<word): grep root <<<"$(id)" feeds one string as stdin.
Control operators: &&, ||, ; and newline
make && make install # install ONLY if build succeeded
make || printf 'build failed\n' >&2
command1; command2 # sequence, ignore outcomes
These are control operators, not pipes: no bytes flow between the commands; only exit STATUS does. && runs the right side only if the left succeeded; || only if it failed (short-circuit); ; and newline sequence unconditionally. Note the >&2; diagnostics belong on stderr even in one-liners.
Exit status: the only message passing between commands
true
echo $? # -> 0
false
echo $? # -> 1
Zero means success by Unix convention. Non-zero means failure, and the SPECIFIC value belongs to the command; grep uses 1 for 'no lines matched' (and 2 for real trouble), curl has its own catalogue. Scripts should treat non-zero as failure and consult the manual of the specific tool before decoding further.
Pipeline status: the default, pipefail, and PIPESTATUS
Verified against the installed bash manual: the return status of a pipeline is the exit status of its LAST command; unless pipefail is set, in which case it is the rightmost failing command's status, or zero if all succeed.
false | true
echo $? # -> 0 (surprise: the failure vanished)
set -o pipefail
false | true
echo $? # -> 1 (the failure is visible again)
Why the default is 0: pipelines are designed for data flow, and consumers like head routinely exit early (see SIGPIPE below). But it means an earlier stage can fail while the pipeline 'succeeds'. To see every stage's status at once, Bash exposes PIPESTATUS:
false | true | grep -q x
echo "${PIPESTATUS[@]}" # -> 1 0 1
# false grep-no-match true... read as:
# stage1=1, stage2=1(grep matched nothing? check), stage3=...
# Practical readout after any pipeline:
for st in "${PIPESTATUS[@]}"; do printf '%d ' "$st"; done; echo
PIPESTATUS is overwritten by the NEXT pipeline or simple command. Read it immediately, and note that echo-ing it first is itself a command; capture it: st=("${PIPESTATUS[@]").
set -euo pipefail: three switches, not an incantation
| Option | What it does | Read this carefully |
|---|---|---|
| -e | Shell exits when a command fails; with important contextual exceptions (commands in && / || lists, conditions in if/while, negated commands, and pipelines whose status is checked do NOT trigger it) | The installed Bash manual's 'set -e' section spells out the exceptions. It is not simply 'fail fast'; it is 'fail on unguarded failures' |
| -u | Expanding an unset variable is an error | Catches the $OWER typo that silently deleted the wrong directory. Also makes "${VAR:-}" the idiom for deliberately-optional values |
| -o pipefail | Pipeline status reflects a failing component | Turns 'earlier stage failed silently' into an error, as shown above |
Use the triple in scripts you intend to keep. Understand it before you put it in scripts other people run: -e's exceptions mean some failures still pass unguarded, and a shell that exits mid-script can leave half-done work; pair it with explicit cleanup via trap where that matters.
# Debugging execution: trace every command before it runs.
set -x
# trace lines look like: + ls -l /etc/network.d
# Controlled trace with a prefix you can grep:
PS4='+ ${BASH_SOURCE}:${LINENO}: ' set -x
Safe filenames: newline is data
The single most common pipeline bug in the wild: treating newline as a filename separator. Newline is a LEGAL character inside filenames, and spaces are everywhere. The classic unsafe loop:
# UNSAFE: filenames with spaces/newlines explode into multiple words,
# and glob characters in names may expand to OTHER files.
for file in $(find . -type f); do
...
done
Verified from the installed manuals: find -print0 terminates each name with a NUL byte, and xargs -0 consumes NUL-separated items. NUL is the one byte no filename may contain, which makes it the only honest separator.
# Safe: NUL-delimited handoff
find . -type f -print0 | xargs -0 grep -l 'TODO'
# Safe without xargs, batched by find itself:
find . -type f -exec grep -l 'TODO' {} +
# Safe shell loop over find:
find . -type f -print0 |
+ while IFS= read -r -d '' file; do
+ printf '%s\n' "$file"
+ done
And the conceptual distinction that ties back to arguments-versus-streams: producer | xargs command converts a STREAM into ARGUMENTS (batched with -n N when lists get long); producer | command hands the stream straight through. Different plumbing, different failure modes, both legitimate.
The toolbox as roles, not commands
| Tool | Pipeline role | One-line example |
|---|---|---|
| grep | select records | grep ERROR app.log |
| cut | select simple fields | cut -d: -f1 /etc/passwd |
| tr | translate/delete characters | tr a-z A-Z |
| sort | order records | sort -rn |
| uniq | collapse ADJACENT duplicates | sort | uniq -c |
| wc | count lines/words/bytes | wc -l |
| sed | transform streams | sed 's/foo/bar/g' |
| awk | parse and compute structured text | awk '{sum+=$2} END{print sum}' |
sed and awk have their own chapter; this page stops at their pipeline roles and hands over there for the deep material. The sort|uniq pairing deserves one more sentence forever: uniq is an ADJACENCY operator. Unsorted input is not deduplicated; it is merely reported accurately about neighbours.
Real administrative pipelines
Each example names its question first. All tools verified installed on Saphira (coreutils, findutils, diffutils, gawk, sed, grep, bash, procps, iproute2).
# Q: What is eating this disk, biggest first?
df -h | sort -k5 -rn | head
# stages: human-readable table -> sort by 5th col (use%) desc -> worst 10
# Q: Which processes hold the most memory right now?
ps aux | sort -rnk4 | head -8
# Q: How many packages are installed, and which came from testing?
apk info | wc -l
apk info -v 2>/dev/null | grep -- '-r' | head # repo-tagged versions
# Q: What is listening, on which ports?
ss -tulnp | awk 'NR>1 {print $1, $5}' | sort | uniq -c | sort -rn
# Q: Who has been hammering the mail log today?
grep 'auth failed' /var/log/messages | awk '{print $NF}' | sort | uniq -c | sort -rn | head
# Q: Top client addresses in an HTTP access log (see the worked
# build below for the full derivation).
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head
# Q: Did my last edit actually change anything?
diff <(sort old.list) <(sort new.list) # Bash process substitution
Notice the shape repeated across all of them: select -> normalise/order -> aggregate -> rank -> limit. That shape is a template you can retarget at any log, any inventory, any list.
Streams, records, and the limits of text
Unix tools treat newline-delimited text as records, but the pipe carries BYTES. A pipeline fed binary data behaves according to each tool's C locale instincts: grep may say 'binary file matches', awk may mangle, wc -l counts NULs happily. Embedded newlines inside CSV fields are indistinguishable from record boundaries to sort and uniq.
Parsing human-formatted command output (ps aux columns, df tables) is a convenience, not a contract: column positions move between versions and locales. When a script matters, prefer a machine-readable source; /proc files, apk --json where offered, SQL, or explicit field joins; over re-reading prose meant for humans.
Locales: when 'sorted' needs a definition
sort and character-class matching depend on locale collation rules. Saphira is musl-based, where locale support is deliberately minimal, but the ordering rules still shift between the C locale and UTF-8 handling; enough to make one machine's script sort differently on another.
# Deterministic scripts pin the locale:
LC_ALL=C sort userlist.txt
# Demonstrate the difference once and you will never forget it:
printf '%s\n' a B c A | sort # locale-aware order
printf '%s\n' a B c A | LC_ALL=C sort # byte order: A B a c
Buffering: why output arrives in bursts
A program writing to a TERMINAL is usually line-buffered (you see each line). The same program writing into a PIPE is usually block-buffered (4-8 KiB accumulates first), so command A | command B can look hung while A quietly fills a buffer. Nothing is wrong; libc chose a different mode because nobody is watching a screen.
# Force line buffering on the producer where the tool offers it:
stdbuf -oL slow_program | consumer # stdbuf: verified installed (coreutils)
# Or make the producer flush explicitly (awk):
awk '{print $1; fflush()}' big.log | tail -f
tail -f on a live file bypasses the problem entirely for log watching. The skill is recognising 'bursts of output' as a buffering artefact, not a bug.
SIGPIPE: when head walks out of the restaurant
something-that-produces-a-lot | head -10
head reads ten lines, prints them, exits. The producer, still writing, now faces a pipe with no reader and receives SIGPIPE (signal 13); normally terminating it. That is the DESIGN working: nobody had to invent a 'stop producing' protocol. If a producer logs 'broken pipe' or exits 141 (128+13), that is usually this healthy mechanism, not a mysterious crash.
yes | head -3
printf '%s\n' "${PIPESTATUS[@]}" # -> 141 0 : yes died by SIGPIPE, head fine
Named pipes: pipelines with a filesystem address
mkfifo /tmp/example.pipe
producer > /tmp/example.pipe & # writer blocks until a reader opens
count < /tmp/example.pipe # reader drains it
An anonymous pipeline exists only inside one shell command line and dies with it. A FIFO is a filesystem object (verify with ls -l: the type is p) that any two unrelated processes can open by name; a pipeline with a doorbell. Same kernel buffer machinery, different lifetime and discoverability.
Process substitution: a file that is really a pipeline
diff <(sort old.txt) <(sort new.txt)
The problem it solves: diff insists on filenames, but you want it to compare two COMMAND outputs. <(...) runs the command, connects its stdout to a fresh pipe/fd, and passes a magic filename like /dev/fd/63. It composes anywhere a file argument is expected. Labelled honestly: this is Bash/process-substitution territory; it does not exist in plain POSIX sh, unlike everything earlier on this page.
Shell grammar is not punctuation soup
Decode one full line, token by token; what the SHELL does before any program begins:
command <input 2>errors | transform | tee output | final
-
1. Parse operators first
The shell splits the line at | into four segments; <, 2>, | are grammar to the shell; they never become arguments.
-
2. Create the pipe(s)
For each |, the shell calls pipe(2): three kernel buffers, six file-descriptor ends.
-
3. Apply redirections per segment
segment 1: fd0 <- input, fd2 -> errors; segment 3: fd1 duplicated to tee's stdout AND 'output'.
-
4. Fork and exec
Each segment becomes a process holding its rewired descriptors. 'command' cannot see that stderr went to a file; 'final' cannot tell its stdin was ever a terminal.
-
5. Wire and run
fd1 of each left segment is connected to fd0 of the next. Then, and only then, does any byte move.
Every symbol in the line is shell grammar acting on descriptors. The programs are bystanders to the plumbing.
Build a pipeline from an English question
Which ten client IP addresses made the most requests to this HTTP access log? Work the sentence into stages before touching the keyboard:
-
1. Need the IP field
It is the first whitespace-separated field in this log format.
-
2. Extract the first field
awk '{print $1}' access.log
-
3. Put equal IPs together
sort
-
4. Count neighbours
uniq -c
-
5. Rank numerically, descending
sort -rn
-
6. Take the first ten
head -10
awk '{print $1}' access.log |
+ sort |
+ uniq -c |
+ sort -rn |
+ head -10
State the assumptions or they will bite: this works only while the first field really is the client address. Behind HAProxy with PROXY protocol, or after an intermediate NAT, field one may be the proxy, and the pipeline will confidently report the wrong truth. Verify field one against a known request before trusting the ranking.
Pipeline troubleshooting: the classics and their cures
| Symptom | Investigate |
|---|---|
| Pipeline produces nothing | Run the first stage alone. Then add stages one at a time; the empty result arrives at a specific boundary. |
| Errors bypass the pipe | They are on stderr by design. Re-plumb with 2>&1 | when the consumer should see them. |
| Command appears hung | It is probably waiting on stdin. Check whether an earlier stage actually connected; a filter with no input reads forever. |
| Output arrives in bursts | Block buffering into a pipe (see the buffering section); stdbuf -oL or an explicit flush. |
| Earlier stage failed, final stage returned success | Default pipeline status is the last command. set -o pipefail, then read "${PIPESTATUS[@]}". |
| Quotes disappeared in output | They were consumed by the shell as syntax. Quote correctly or use here-docs with a quoted delimiter. |
| Filenames with spaces broke the loop | $(find) word-splitting. -print0 | xargs -0, find -exec ... +, or read -d ''. |
| Output file was unexpectedly emptied | > truncates BEFORE the command runs; a failed producer still empties the target. Use >| only deliberately; write to temp then mv. |
| sudo command >file fails to write | The SHELL performs redirection as you, before sudo starts. sudo sh -c 'command >file', or tee: sudo command | sudo tee file. |
| Producer died after head exited | SIGPIPE by design (exit 141). Not a crash; stop worrying unless the producer's partial state matters. |
| Works interactively, fails in cron/CI | Environment differs: PATH, locale, terminal buffering, cwd. Pin what the script assumes (PATH, LC_ALL=C, absolute paths). |
Exercises
Attempt each before opening the answer. The progression is deliberate: beginner to administrator.
Beginner: run a command whose errors and results you want SEPARATELY; send stdout to results.txt, stderr to problems.txt, without losing either.
Beginner: make a pipeline that reports HOW MANY packages are installed and prints the five most recently modified files under /etc, safest form.
Intermediate: find the ten largest files under /srv safely (any filename), largest first, showing size and path.
Intermediate: prove which stage of 'grep ERROR app.log | filter1 | sort | uniq -c' is failing when the final output is empty, without touching the command's code.
Advanced: extract the client IP field from an HTTP access log into ips.txt AND keep any 'permission denied' complaints visible on your terminal, in one command.
Administrator: a five-stage pipeline works interactively but produces partial output in cron. Diagnose and fix without changing the stages.
Did we miss something?
If this page left something unanswered, found an error, or there is another subject you would like documented, tell us. Saphira’s documentation grows from real problems people need to solve.
Send feedback or request a new section →
Prefer not to do it yourself?
Everything needed to do the work yourself is documented here and remains free; we charge for human time, not for withholding knowledge. Sometimes the missing resource is simply time. The same people who build Saphira can provide paid professional help with implementation, migration, troubleshooting and administration.