JEPA4Japan · tutorials

Strings and Text Processing

1,006 words 5 min read #Python

Slice, search, normalize, and format text without losing sight of Unicode.

Course progress Course outline 24 of 24 lessons available

Text is a row of tiny tiles

A Python string is text in quotes: "Python", "こんにちは", or "👍🏽". Imagine its pieces sitting in a row. Python calls those pieces Unicode code points.

  1. Get text a name, note, or tag
  2. Clean it spaces and spelling style
  3. Search it find useful pieces
  4. Make new text the old string stays safe
Strings let Python inspect and rebuild text.

Python counts positions from 0. len() counts code points, not bytes and not always the symbols a person sees.

word = "Python"

print(len(word))
print(word[0])
print(word[-1])
print(word[1:4])
6
P
n
yth

word[0] is the first tile. word[-1] is the last. A slice includes its start but stops before its end, so [1:4] takes positions 1, 2, and 3.

Point, slice, split, and join

A single index must exist. word[99] raises IndexError. A slice is gentler: word[:99] simply returns all of word. An empty string has no first or last position, so check it before indexing:

name = "".strip()

if name:
    print(name[0])
else:
    print("No first character")

split() cuts text into a list. join() glues string pieces together with a chosen separator:

messy = "  read   code\tplay  "
pieces = messy.split()

print(pieces)
print(" ".join(pieces))
['read', 'code', 'play']
read code play
  1. Split text becomes pieces
  2. List pieces wait in order
  3. Join pieces become text
" ".join(text.split()) collapses accidental whitespace.

With split(","), commas are the exact dividers. Two touching commas create an empty piece. Also, every item passed to join() must be a string.

Cleaners return a new string

Strings are immutable: Python cannot replace one tile inside an existing string. This is an intentional error:

word = "python"
word[0] = "P"

It raises TypeError. Build a new string instead:

word = "python"
title = "P" + word[1:]

print(word)
print(title)
python
Python

String methods follow the same rule. strip() removes surrounding whitespace, replace() swaps literal pieces, and casefold() makes a strong Unicode-aware search key. Each returns new text.

raw = "  Python makes TEXT  "
clean = raw.strip().replace("makes", "cleans")
searchable = clean.casefold()

print(clean)
print("python" in searchable)
print(searchable.count("text"))
print(searchable.find("cleans"))
Python cleans TEXT
True
1
7

Use in for “is it there?”, count() for “how many non-overlapping matches?”, and find() for the first position (-1 means missing). Never use if text.find(term):: a match at position 0 acts false, while -1 acts true.

One picture can use several code points

Two strings can look identical yet contain different code-point rows. Python’s standard-library unicodedata module can put canonically equivalent text into the same Unicode form.

  1. One tile a composed letter
  2. Two tiles a letter plus an accent
  3. NFC makes equivalent rows match
Same appearance does not always mean the same code points.
import unicodedata

one = "café"
two = "cafe\u0301"

print(one == two)
print(len(one), len(two))
print(one == unicodedata.normalize("NFC", two))
print(len("👍🏽"))
False
4 5
True
2

The emoji looks like one symbol but uses two code points. Basic indexing can split it. String positions are therefore neither byte positions nor reliable “visible character” boundaries.

Tiny project: Study Note Cleaner

Create note_cleaner.py. This complete program uses only the standard library. It normalizes Unicode, collapses whitespace, and counts a non-empty keyword without changing the original strings.

import unicodedata


def search_key(text):
    return unicodedata.normalize("NFC", text).strip().casefold()


def clean_note(text):
    normalized = unicodedata.normalize("NFC", text)
    return " ".join(normalized.split())


raw_note = "  CAFÉ   practice, then cafe\u0301 review.  "
raw_keyword = " Café "

note = clean_note(raw_note)
keyword = search_key(raw_keyword)
matches = search_key(note).count(keyword)

print(f"Original: <{raw_note}>")
print(f"Clean: <{note}>")
print(f"Matches: {matches}")
print(f"Original unchanged: {raw_note.startswith('  ')}")

Run python3 note_cleaner.py and check:

Original: <  CAFÉ   practice, then café review.  >
Clean: <CAFÉ practice, then café review.>
Matches: 2
Original unchanged: True

Boundary rule: the keyword must not be empty. text.count("") counts gaps and returns len(text) + 1, which is not a useful study-word count.

Three tiny missions

  1. Safe initials. Strip an input word. Print its first and last code points only when it is not empty.
  2. Tiny preview. Write preview(text, limit) using text[:limit]. Test 0, 3, and a limit bigger than the text.
  3. Search detective. Find a term at position 0 and a missing term. Use in or compare find() with -1 correctly.

You are ready for Chapter 6 when…

  • you can use len(), positive and negative indexes, and slices;
  • you can explain that strings are immutable;
  • you can clean text with strip(), split(), join(), and casefold();
  • you can choose in, count(), or find() for the question you mean;
  • you can normalize equivalent Unicode text with unicodedata.normalize();
  • you remember that code points, visible symbols, and bytes are different things;
  • you can run the Note Cleaner and match all four output checks.

Next, you will put the pieces made by split() into lists, then compare changeable lists with fixed tuples.