Skip to content

RTFM · Tools

Shell pipelines: how Unix thinks

This chapter is not six funny symbols. It is the mental model underneath them: small programs that move bytes, connected by the shell, until bytes become a useful answer. Once that model is real to you, grep, sed, awk, sort, tee, find, xargs, logs, and half of Linux administration, stop being mysterious.

Saphira Linux dragon mascot

Every process has three streams

When a program runs, the kernel hands it three open file descriptors before main() executes a line. They are not magic: a file descriptor is just a small integer indexing the process's open-file table, and almost anything can sit behind it; a terminal, a regular file, a pipe, a socket, or /dev/null.

The three descriptors
                       process
                   +-------------+
keyboard/file ---->| 0  stdin    |
                   |             |----> 1 stdout
                   |             |
                   |             |----> 2 stderr
                   +-------------+
The cast
DescriptorNameDefault connectionIntended cargo
0stdinkeyboard (terminal)data to consume
1stdoutterminalthe work product: results, logs-as-data
2stderrterminaldiagnostics: errors, warnings, progress

The separation of 1 and 2 is the whole trick. A well-behaved Unix tool puts RESULTS on stdout and NOISE on stderr, so results can be piped onward without diagnostics contaminating them. Commands that blur this deserve your suspicion.

Redirection is re-plumbing descriptors

Every redirection below is the shell rewiring one of those three integers before the command starts. The command itself cannot tell a terminal from a file, and that is the point.

The redirection vocabulary
command >output.txt              # fd1 -> file (truncate)
command >>output.txt             # fd1 -> file (append)
command 2>errors.txt             # fd2 -> file
command >output.txt 2>errors.txt # fd1 and fd2 -> separate files
command >everything.txt 2>&1     # fd1 -> file, THEN fd2 -> wherever fd1 points
command 2>&1 >output.txt         # fd2 -> terminal(!), then fd1 -> file
command </path/to/input          # fd0 -> file
command >/dev/null               # discard results
command 2>/dev/null              # discard diagnostics

The two 2>&1 forms are where everybody trips, so compare them side by side. 2>&1 means 'duplicate descriptor 2 onto descriptor 1; wherever fd1 points RIGHT NOW'. It is a copy of a pointer, taken at that moment, not a permanent linkage.

Same words, different order, opposite outcome
command >everything.txt 2>&1

# step 1: fd1 --> everything.txt
# step 2: fd2 --> everything.txt   (fd1 already points there)
#
# result: both streams land in the file. Correct.

command 2>&1 >output.txt

# step 1: fd2 --> (wherever fd1 points NOW: the terminal)
# step 2: fd1 --> output.txt
#
# result: errors stay on screen, results go to the file.
# The streams were crossed in the wrong order.

Bash offers >& as in command >&file and &>file for the common both-to-file case, but the numbered form above is portable and teaches the mechanism. Learn the mechanism once and the shorthand reads itself.

Pipes: stdout into stdin, nothing in between

The only picture that matters
stdout                         stdin
command A ------- pipe -------> command B

The | operator connects fd1 of the left command to fd0 of the right command; a kernel buffer in the middle, not a file. Two important truths follow. First, the shell creates the pipe, forks both processes, and hands each its end; the data never passes through any temporary file, which is why streaming a 40 GB log works in 4 GB of RAM. Second, stderr is NOT piped: command A's diagnostics still write to your terminal, by design, so errors stay visible while data flows.

Piping stderr too: three ways, one habit

Three forms, verified against the installed bash manual
command | consumer             # data only; errors stay on terminal
command 2>&1 | consumer        # data + errors, explicitly re-plumbed
command |& consumer            # Bash shorthand for the line above

Use plain | when you want errors visible and data flowing. Use |& (or the explicit 2>&1 |) when the consumer should also see the diagnostics; for example when feeding a filter that selects ERROR lines from a noisy build.

|& is a Bash extension, not POSIX sh. It is correct on Saphira (bash ships), but scripts claiming portability should write 2>&1 | explicitly.

First pipelines, stage by stage

Count three words:

One pipe
printf '%s\n' alpha beta gamma | wc -l
# -> 3

printf emits three lines (the %s\n format is applied per argument), wc -l counts newlines. Now a two-stage pipeline where every intermediate result is shown - this is the habit that makes pipelines legible:

Every stage visible
printf '%s\n' banana apple banana pear | sort
# stage 1 output:
#   apple
#   banana
#   banana
#   pear

printf '%s\n' banana apple banana pear | sort | uniq -c
# stage 2 output:
#      1 apple
#      2 banana
#      1 pear

Read it again with the model in mind: bytes left to right, each program transforming a stream it knows nothing about. Nobody parsed anything; uniq merely noticed that equal lines are now ADJACENT, which is why sort | uniq is a pair, not a coincidence. uniq collapses adjacent duplicates only; fed unsorted input it reports duplicates that are not there.

tee: watch a pipeline without dismantling it

Inspection taps
command1 |
+  tee /tmp/stage1 |
+  command2 |
+  tee /tmp/stage2 |
+  command3

tee copies its stdin to a file AND to its stdout, so you can inspect exactly what flowed between stages after the fact, while the pipeline runs untouched. When a five-stage pipeline misbehaves, tee turns 'something in here is wrong' into 'stage 2 receives garbage', which is a different, tractable problem.

Quoting and expansion: the shell edits your words first

Before any command runs, the shell performs expansion on its arguments: parameter expansion, globbing, command substitution, word splitting. The command never sees your quotes; they are instructions to the shell, consumed during parsing.

The expansion cast
"$variable"      # expands, then stays ONE word even if it contains spaces
'$literal'       # nothing expands; dollars and stars stay literal
*                # expands to matching filenames in the current directory
?                # expands to any single character
$(command)       # runs command, substitutes its stdout

Why quoting matters, demonstrated by the difference between two lines that look like twins:

Not equivalent, ever
file="my notes.txt; rm -rf ~/docs"

rm $file     # shell splits on spaces -> rm my notes.txt; rm -rf ~/docs  (disaster)
rm "$file"   # one word, exactly the filename                     (correct)

Also keep two different kinds of 'input' separate in your head: TEXT FLOWING THROUGH A PIPE (streaming bytes between already-running processes) and COMMAND-LINE ARGUMENTS (a fixed list fixed before the program starts, subject to all the expansion rules above). sort <file reads a stream; sort file1 file2 takes arguments. Pipes carry streams; arguments are built once.

Grouping: { } versus ( )

Two kinds of grouping
{ command1; command2; } >combined.log

(
    cd /somewhere
    command
)

Braces { ...; } run the commands in the CURRENT shell; variables set inside survive, and the redirection applies to the group. Parentheses ( ... ) create a subshell: a child copy of the shell whose cd, variable changes and umask vanish when it exits. Use ( ) precisely when you want throwaway environment; the cd above cannot leak out and disturb the caller.

Command substitution

Capture output into a variable
kernel="$(uname -r)"
printf 'Running kernel: %s\n' "$kernel"

$(...) is preferred over the old backtick syntax: it nests cleanly, reads cleanly, and its failure modes are quieter. When the substitution produces multiple words, the shell splits them; quote "$(...)" when you want one word, exactly as with variables.

Here-documents and here-strings

Quoted versus unquoted delimiter
cat <<'EOF'
literal text
$nothing expands here
EOF

cat <<EOF
the kernel is $(uname -r)   # expansion happens
EOF

Quoting the delimiter ('EOF') turns expansion OFF for the whole block; leaving it bare turns expansion ON. This distinction matters hugely the day you use a heredoc to generate a config file: one form writes literal $ variables for the target program, the other bakes in the generating shell's values. Bash also offers here-strings (verified in the local manual as [n]<<<word): grep root <<<"$(id)" feeds one string as stdin.

Control operators: &&, ||, ; and newline

Sequencing with intent
make && make install        # install ONLY if build succeeded
make || printf 'build failed\n' >&2
command1; command2           # sequence, ignore outcomes

These are control operators, not pipes: no bytes flow between the commands; only exit STATUS does. && runs the right side only if the left succeeded; || only if it failed (short-circuit); ; and newline sequence unconditionally. Note the >&2; diagnostics belong on stderr even in one-liners.

Exit status: the only message passing between commands

Zero is success
true
echo $?    # -> 0

false
echo $?    # -> 1

Zero means success by Unix convention. Non-zero means failure, and the SPECIFIC value belongs to the command; grep uses 1 for 'no lines matched' (and 2 for real trouble), curl has its own catalogue. Scripts should treat non-zero as failure and consult the manual of the specific tool before decoding further.

Pipeline status: the default, pipefail, and PIPESTATUS

Verified against the installed bash manual: the return status of a pipeline is the exit status of its LAST command; unless pipefail is set, in which case it is the rightmost failing command's status, or zero if all succeed.

What pipefail changes
false | true
echo $?          # -> 0   (surprise: the failure vanished)

set -o pipefail
false | true
echo $?          # -> 1   (the failure is visible again)

Why the default is 0: pipelines are designed for data flow, and consumers like head routinely exit early (see SIGPIPE below). But it means an earlier stage can fail while the pipeline 'succeeds'. To see every stage's status at once, Bash exposes PIPESTATUS:

Inspect every stage
false | true | grep -q x
echo "${PIPESTATUS[@]}"    # -> 1 0 1
#                             false grep-no-match true... read as:
#                             stage1=1, stage2=1(grep matched nothing? check), stage3=...

# Practical readout after any pipeline:
for st in "${PIPESTATUS[@]}"; do printf '%d ' "$st"; done; echo

PIPESTATUS is overwritten by the NEXT pipeline or simple command. Read it immediately, and note that echo-ing it first is itself a command; capture it: st=("${PIPESTATUS[@]").

set -euo pipefail: three switches, not an incantation

Each switch explained
OptionWhat it doesRead this carefully
-eShell exits when a command fails; with important contextual exceptions (commands in && / || lists, conditions in if/while, negated commands, and pipelines whose status is checked do NOT trigger it)The installed Bash manual's 'set -e' section spells out the exceptions. It is not simply 'fail fast'; it is 'fail on unguarded failures'
-uExpanding an unset variable is an errorCatches the $OWER typo that silently deleted the wrong directory. Also makes "${VAR:-}" the idiom for deliberately-optional values
-o pipefailPipeline status reflects a failing componentTurns 'earlier stage failed silently' into an error, as shown above

Use the triple in scripts you intend to keep. Understand it before you put it in scripts other people run: -e's exceptions mean some failures still pass unguarded, and a shell that exits mid-script can leave half-done work; pair it with explicit cleanup via trap where that matters.

set -x and PS4
# Debugging execution: trace every command before it runs.
set -x
# trace lines look like:  + ls -l /etc/network.d

# Controlled trace with a prefix you can grep:
PS4='+ ${BASH_SOURCE}:${LINENO}: ' set -x

Safe filenames: newline is data

The single most common pipeline bug in the wild: treating newline as a filename separator. Newline is a LEGAL character inside filenames, and spaces are everywhere. The classic unsafe loop:

Do not write this
# UNSAFE: filenames with spaces/newlines explode into multiple words,
# and glob characters in names may expand to OTHER files.
for file in $(find . -type f); do
    ...
done

Verified from the installed manuals: find -print0 terminates each name with a NUL byte, and xargs -0 consumes NUL-separated items. NUL is the one byte no filename may contain, which makes it the only honest separator.

Three correct forms
# Safe: NUL-delimited handoff
find . -type f -print0 | xargs -0 grep -l 'TODO'

# Safe without xargs, batched by find itself:
find . -type f -exec grep -l 'TODO' {} +

# Safe shell loop over find:
find . -type f -print0 |
+  while IFS= read -r -d '' file; do
+      printf '%s\n' "$file"
+  done

And the conceptual distinction that ties back to arguments-versus-streams: producer | xargs command converts a STREAM into ARGUMENTS (batched with -n N when lists get long); producer | command hands the stream straight through. Different plumbing, different failure modes, both legitimate.

The toolbox as roles, not commands

One job each
ToolPipeline roleOne-line example
grepselect recordsgrep ERROR app.log
cutselect simple fieldscut -d: -f1 /etc/passwd
trtranslate/delete characterstr a-z A-Z
sortorder recordssort -rn
uniqcollapse ADJACENT duplicatessort | uniq -c
wccount lines/words/byteswc -l
sedtransform streamssed 's/foo/bar/g'
awkparse and compute structured textawk '{sum+=$2} END{print sum}'

sed and awk have their own chapter; this page stops at their pipeline roles and hands over there for the deep material. The sort|uniq pairing deserves one more sentence forever: uniq is an ADJACENCY operator. Unsorted input is not deduplicated; it is merely reported accurately about neighbours.

Real administrative pipelines

Each example names its question first. All tools verified installed on Saphira (coreutils, findutils, diffutils, gawk, sed, grep, bash, procps, iproute2).

System questions
# Q: What is eating this disk, biggest first?
df -h | sort -k5 -rn | head
# stages: human-readable table -> sort by 5th col (use%) desc -> worst 10

# Q: Which processes hold the most memory right now?
ps aux | sort -rnk4 | head -8

# Q: How many packages are installed, and which came from testing?
apk info | wc -l
apk info -v 2>/dev/null | grep -- '-r' | head    # repo-tagged versions

# Q: What is listening, on which ports?
ss -tulnp | awk 'NR>1 {print $1, $5}' | sort | uniq -c | sort -rn
Log and diff questions
# Q: Who has been hammering the mail log today?
grep 'auth failed' /var/log/messages | awk '{print $NF}' | sort | uniq -c | sort -rn | head

# Q: Top client addresses in an HTTP access log (see the worked
#    build below for the full derivation).
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head

# Q: Did my last edit actually change anything?
diff <(sort old.list) <(sort new.list)     # Bash process substitution

Notice the shape repeated across all of them: select -> normalise/order -> aggregate -> rank -> limit. That shape is a template you can retarget at any log, any inventory, any list.

Streams, records, and the limits of text

Unix tools treat newline-delimited text as records, but the pipe carries BYTES. A pipeline fed binary data behaves according to each tool's C locale instincts: grep may say 'binary file matches', awk may mangle, wc -l counts NULs happily. Embedded newlines inside CSV fields are indistinguishable from record boundaries to sort and uniq.

Parsing human-formatted command output (ps aux columns, df tables) is a convenience, not a contract: column positions move between versions and locales. When a script matters, prefer a machine-readable source; /proc files, apk --json where offered, SQL, or explicit field joins; over re-reading prose meant for humans.

Locales: when 'sorted' needs a definition

sort and character-class matching depend on locale collation rules. Saphira is musl-based, where locale support is deliberately minimal, but the ordering rules still shift between the C locale and UTF-8 handling; enough to make one machine's script sort differently on another.

Pin it when it matters
# Deterministic scripts pin the locale:
LC_ALL=C sort userlist.txt

# Demonstrate the difference once and you will never forget it:
printf '%s\n' a B c A | sort          # locale-aware order
printf '%s\n' a B c A | LC_ALL=C sort # byte order: A B a c

Buffering: why output arrives in bursts

A program writing to a TERMINAL is usually line-buffered (you see each line). The same program writing into a PIPE is usually block-buffered (4-8 KiB accumulates first), so command A | command B can look hung while A quietly fills a buffer. Nothing is wrong; libc chose a different mode because nobody is watching a screen.

Practical controls
# Force line buffering on the producer where the tool offers it:
stdbuf -oL slow_program | consumer     # stdbuf: verified installed (coreutils)

# Or make the producer flush explicitly (awk):
awk '{print $1; fflush()}' big.log | tail -f

tail -f on a live file bypasses the problem entirely for log watching. The skill is recognising 'bursts of output' as a buffering artefact, not a bug.

SIGPIPE: when head walks out of the restaurant

The everyday SIGPIPE
something-that-produces-a-lot | head -10

head reads ten lines, prints them, exits. The producer, still writing, now faces a pipe with no reader and receives SIGPIPE (signal 13); normally terminating it. That is the DESIGN working: nobody had to invent a 'stop producing' protocol. If a producer logs 'broken pipe' or exits 141 (128+13), that is usually this healthy mechanism, not a mysterious crash.

Watch it happen
yes | head -3
printf '%s\n' "${PIPESTATUS[@]}"   # -> 141 0 : yes died by SIGPIPE, head fine

Named pipes: pipelines with a filesystem address

A FIFO in three moves
mkfifo /tmp/example.pipe
producer > /tmp/example.pipe &   # writer blocks until a reader opens
count < /tmp/example.pipe        # reader drains it

An anonymous pipeline exists only inside one shell command line and dies with it. A FIFO is a filesystem object (verify with ls -l: the type is p) that any two unrelated processes can open by name; a pipeline with a doorbell. Same kernel buffer machinery, different lifetime and discoverability.

Process substitution: a file that is really a pipeline

Bash-specific (not POSIX sh)
diff <(sort old.txt) <(sort new.txt)

The problem it solves: diff insists on filenames, but you want it to compare two COMMAND outputs. <(...) runs the command, connects its stdout to a fresh pipe/fd, and passes a magic filename like /dev/fd/63. It composes anywhere a file argument is expected. Labelled honestly: this is Bash/process-substitution territory; it does not exist in plain POSIX sh, unlike everything earlier on this page.

Shell grammar is not punctuation soup

Decode one full line, token by token; what the SHELL does before any program begins:

The specimen
command <input 2>errors | transform | tee output | final
  1. 1. Parse operators first

    The shell splits the line at | into four segments; <, 2>, | are grammar to the shell; they never become arguments.

  2. 2. Create the pipe(s)

    For each |, the shell calls pipe(2): three kernel buffers, six file-descriptor ends.

  3. 3. Apply redirections per segment

    segment 1: fd0 <- input, fd2 -> errors; segment 3: fd1 duplicated to tee's stdout AND 'output'.

  4. 4. Fork and exec

    Each segment becomes a process holding its rewired descriptors. 'command' cannot see that stderr went to a file; 'final' cannot tell its stdin was ever a terminal.

  5. 5. Wire and run

    fd1 of each left segment is connected to fd0 of the next. Then, and only then, does any byte move.

Every symbol in the line is shell grammar acting on descriptors. The programs are bystanders to the plumbing.

Build a pipeline from an English question

Which ten client IP addresses made the most requests to this HTTP access log? Work the sentence into stages before touching the keyboard:

  1. 1. Need the IP field

    It is the first whitespace-separated field in this log format.

  2. 2. Extract the first field

    awk '{print $1}' access.log

  3. 3. Put equal IPs together

    sort

  4. 4. Count neighbours

    uniq -c

  5. 5. Rank numerically, descending

    sort -rn

  6. 6. Take the first ten

    head -10

The derived answer
awk '{print $1}' access.log |
+  sort |
+  uniq -c |
+  sort -rn |
+  head -10

State the assumptions or they will bite: this works only while the first field really is the client address. Behind HAProxy with PROXY protocol, or after an intermediate NAT, field one may be the proxy, and the pipeline will confidently report the wrong truth. Verify field one against a known request before trusting the ranking.

Pipeline troubleshooting: the classics and their cures

Symptom -> investigation
SymptomInvestigate
Pipeline produces nothingRun the first stage alone. Then add stages one at a time; the empty result arrives at a specific boundary.
Errors bypass the pipeThey are on stderr by design. Re-plumb with 2>&1 | when the consumer should see them.
Command appears hungIt is probably waiting on stdin. Check whether an earlier stage actually connected; a filter with no input reads forever.
Output arrives in burstsBlock buffering into a pipe (see the buffering section); stdbuf -oL or an explicit flush.
Earlier stage failed, final stage returned successDefault pipeline status is the last command. set -o pipefail, then read "${PIPESTATUS[@]}".
Quotes disappeared in outputThey were consumed by the shell as syntax. Quote correctly or use here-docs with a quoted delimiter.
Filenames with spaces broke the loop$(find) word-splitting. -print0 | xargs -0, find -exec ... +, or read -d ''.
Output file was unexpectedly emptied> truncates BEFORE the command runs; a failed producer still empties the target. Use >| only deliberately; write to temp then mv.
sudo command >file fails to writeThe SHELL performs redirection as you, before sudo starts. sudo sh -c 'command >file', or tee: sudo command | sudo tee file.
Producer died after head exitedSIGPIPE by design (exit 141). Not a crash; stop worrying unless the producer's partial state matters.
Works interactively, fails in cron/CIEnvironment differs: PATH, locale, terminal buffering, cwd. Pin what the script assumes (PATH, LC_ALL=C, absolute paths).

Exercises

Attempt each before opening the answer. The progression is deliberate: beginner to administrator.

Beginner: run a command whose errors and results you want SEPARATELY; send stdout to results.txt, stderr to problems.txt, without losing either.
command >results.txt 2>problems.txt # Verify: wc -l results.txt problems.txt # Both files exist with distinct contents because fd1 and fd2 # were re-plumbed to different files before the command started.
Beginner: make a pipeline that reports HOW MANY packages are installed and prints the five most recently modified files under /etc, safest form.
apk info | wc -l find /etc -type f -printf '%T@ %p\n' 2>/dev/null | sort -rn | head -5 | cut -d' ' -f2- # Or the simpler, still-safe variant: find /etc -type f -print0 2>/dev/null | xargs -0 ls -t | head -5 # Note -print0 | xargs -0: /etc contains filenames with spaces.
Intermediate: find the ten largest files under /srv safely (any filename), largest first, showing size and path.
find /srv -type f -print0 | + xargs -0 du -b 2>/dev/null | + sort -rn | head -10 # Why not: du -a /srv | sort -rn | head # That works too, du emits its own two-column format, but the # -print0|xargs -0 shape generalises to tools that take filename # arguments, and survives newlines in names.
Intermediate: prove which stage of 'grep ERROR app.log | filter1 | sort | uniq -c' is failing when the final output is empty, without touching the command's code.
set -o pipefail ... | sort | uniq -c; echo "${PIPESTATUS[@]}" # The array names the failing stage immediately. If all are zero, # the failure is semantic (nothing matched ERROR) rather than # mechanical: re-run stage one alone and inspect its output. # tee taps between stages give the same answer with data instead # of status codes.
Advanced: extract the client IP field from an HTTP access log into ips.txt AND keep any 'permission denied' complaints visible on your terminal, in one command.
awk '{print $1}' access.log 2> >(logger -t logextract) > ips.txt # Process substitution (Bash-specific): stderr goes to a logger # process, stdout to the file. A more portable equivalent: awk '{print $1}' access.log 2>>extract.errors > ips.txt # ...with extract.errors reviewed after. The portable form is # what scripts should use; the >( ) form is for interactive power.
Administrator: a five-stage pipeline works interactively but produces partial output in cron. Diagnose and fix without changing the stages.
Check the classics in order: 1. Locale: add LC_ALL=C before the sort stages - collation differences change uniq adjacency and therefore results. 2. Buffering: cron output is block-buffered into a pipe; stdbuf -oL on the producer, or accept whole-run output. 3. Environment: pin PATH and absolute paths - cron's PATH is not your login PATH; a stage may be silently failing to exec. 4. Exit codes: set -o pipefail and capture "${PIPESTATUS[@]}" into the job log; interactive success may have been the last stage exiting 0 while an upstream stage died. 5. cwd and relative paths: cron starts elsewhere; pin with cd /. Each fix addresses one layer of 'interactive vs non-interactive' divergence; apply them and the two environments converge.

Did we miss something?

If this page left something unanswered, found an error, or there is another subject you would like documented, tell us. Saphira’s documentation grows from real problems people need to solve.

Send feedback or request a new section →

Prefer not to do it yourself?

Everything needed to do the work yourself is documented here and remains free; we charge for human time, not for withholding knowledge. Sometimes the missing resource is simply time. The same people who build Saphira can provide paid professional help with implementation, migration, troubleshooting and administration.

Ask about professional support →