Computer Science › Compilers & Languages
Lexer / Tokenizer
Splitting source text into tokens.
Also known as: lexer, tokenizer, lexical analysis
A lexer (tokenizer) is the first stage of most language processing: it reads raw text and groups characters into tokens — the meaningful units a parser then arranges. It discards what doesn’t matter (whitespace, comments, in most languages) and classifies what does: a number, a name, an operator, a keyword, a string.
source: "x = 1 + 2"
tokens: NAME(x) ASSIGN(=) NUMBER(1) PLUS(+) NUMBER(2)
The lexer handles the fiddly, local rules: where a token starts and ends, that 123 is one number, that "hello world" is a single string (with quotes, not five tokens and a space), that // comment runs to end of line. Many lexers are essentially a set of regular expressions matched in order, either hand-written or generated.
Splitting lexing from parsing keeps both simple: the lexer deals with characters, the parser deals with tokens and structure. That separation is why a parser can be written in terms of “an identifier followed by =” without worrying about whitespace.
The classic mistakes:
- Confusing lexing with parsing. The lexer produces tokens; it doesn’t understand nesting or grammar.
(a + b) * cis still just tokens to a lexer — structure is the parser’s job. - Trying to do structure in the lexer. Tracking bracket depth or expression shape in the lexer tangles the stages and usually belongs in the parser.
- Forgetting string and comment edge cases. Quotes, escapes, multi-line strings and comments-to-end-of-line are where naive tokenizers break. Handle them explicitly.
- Vague positions. Every token should carry its line/column, or error messages can’t point at the problem — a big usability difference.
- Assuming it’s always needed. Tiny formats might skip a formal lexer, but any real language benefits from the clean split.
The lexer is the front door of a language toolchain: it feeds the parser, which feeds the AST, which feeds the interpreter or compiler. It’s small, but it’s where text first becomes structure — and where the quality of error positions is decided.