Skip to content

RTFM · Tools

sed and awk: the classic pair

Two complementary stream editors: sed rewires text line-wise by substitution; awk treats fields as structured data. Together they replace most ad-hoc scripting temptations.

Saphira Linux dragon mascot

sed: substitute, delete, print selectively

Saphira ships sed supporting extended regular expressions via -E, plus the classic substitution and filtering verbs.

sed repertoire
sed 's/old/new/' file           # first occurrence per line
sed 's/old/new/g' file          # every occurrence
sed -E 's/(\w+)@(\w+)/\2@\1/g'  # extended regex + capture swap
sed -n '/ERROR/p' app.log       # silent mode + explicit print
sed '/^#/d;/^$/d' config        # strip comments and blanks
sed -i.bak 's/a/b/g' file       # in-place with backup safety

-i semantics vary between sed flavours regarding backup suffix handling; pass .bak explicitly or verify behaviour on a scratch copy first.

awk: fields, patterns, arithmetic

awk splits each line on whitespace into $1..$NF; the entire line is $0. That model alone covers most log-reduction tasks.

awk patterns
awk '{print $1}' access.log                    # first column
awk '$9 == 500 {print $7}' access.log          # conditionally filter
awk -F: '{print $1, $3}' /etc/passwd           # colon-delimited
awk '/error/ {count++} END {print count}' app.log  # tally by pattern
awk '{sum+=$2} END {printf "%.1f MB\n", sum/1024}' sizes.txt

The END block runs after input ends; perfect place for totals. BEGIN handles pre-input setup. Between them lies the per-record body, conditionally gated by pattern expressions.

Choosing between them

sed vs awk selection heuristics
Task shapeReach for
Search-replace textual patternsed
Extract/filter columns numericallyawk
Multi-step conditional transformationsawk
Simple global character cleanupsed
Aggregations over delimited dataawk

Learning both pays compound interest: outputs from one feed inputs of the other, chained through pipelines endlessly.

Prove it works: stream editor confidence

a one-liner written from scratch replaces a manual spreadsheet edit on a real log within five minutes.

sed addresses, ranges and the -n discipline

Beyond global substitution, sed commands can be gated by ADDRESSES, line numbers or patterns, so an edit applies exactly where intended.

Addresses gate commands
sed '5d' file                 # delete exactly line 5
sed '10,20s/^/# /' file       # comment out lines 10-20 only
sed '/^\[section\]/,/\[/!b' file # (ranges by pattern are legal)
sed -n '/^Aug 27/p' log        # -n silences default printing;
                               # only matching lines are printed
sed '$!d' file                 # address $ = last line; '!d' = delete others

The -n + explicit p combination is sed's filtering mode and pairs beautifully with pipelines: it becomes 'print only where pattern', one step above grep when you also need a transform in the same pass.

A sed address range /start/,/end/ matches EVERY start after its end, not just the first block. For document chunks with repeated markers, use a limit: /start/,/end/{...} with an explicit exit, or prefer awk stateful parsing.

awk as a small program: BEGIN, body, END, arrays

Every awk program has the same anatomy, and once you see it, awk stops being incantations:

The three-part anatomy
awk '
  BEGIN { print "report start" }     # runs once, before input
  $3 > 100 { count++; sum += $3 }     # body: per record, when pattern true
  END {                               # runs once, after input
    printf "%d records over threshold, total %d\n", count, sum
  }
' metrics.txt

Arrays make awk a real aggregation engine; the 'for each key, tally' pattern that underlies half of log analysis:

Count per key, then rank
awk '{ users[$1]++ } END { for (u in users) printf "%5d %s\n", users[u], u }' \
+  /var/log/messages | sort -rn

gawk (verified installed on Saphira) adds conveniences on top of POSIX awk; length(array), sorted_in array sorting, process substitutions, but the anatomy above is portable and is all most administrative one-liners ever need. Fields: -F: changes the split character; $NF is the last field; NR the record number; FS/OFS control input/output separators.

Did we miss something?

If this page left something unanswered, found an error, or there is another subject you would like documented, tell us. Saphira’s documentation grows from real problems people need to solve.

Send feedback or request a new section →

Prefer not to do it yourself?

Everything needed to do the work yourself is documented here and remains free; we charge for human time, not for withholding knowledge. Sometimes the missing resource is simply time. The same people who build Saphira can provide paid professional help with implementation, migration, troubleshooting and administration.

Ask about professional support →