2 February 2026

Sequence mining without drowning in n-grams

Bar chart printouts on a table

Give an analyst an unlimited n-gram extractor and you will receive a PDF nobody reads. Usage pattern analysis needs sequences. It does not need every permutation of settings → billing → settings → tooltip.

Where n-grams earn their keep

Bigrams and short trigrams are useful when your event contract is already tight. They show the step after invite, the return into the approval queue, the leap from empty state to first record. In the Sequence Atlas lab we allow them only after the contract page is signed off. Otherwise you are mining typos.

The junk drawer

Once n grows, the catalogue fills with campaign debris, double-taps, and late SDK batches. Teams then “cluster” the mess and present a slide called journeys. That slide cannot survive a hallway question. We teach a hard cap: twenty named paths in the living atlas. Everything else is appendix or deletion.

Collapsing is a judgement, not a script. Faculty will argue with you about whether “export CSV” and “download report” are the same step. That argument is the work. An algorithm that merges them for you is just hiding the fight.

Tool-agnostic on purpose

We do not teach a particular miner. SQL window functions, a notebook, or a printed sticky wall can all produce the same twenty cards. What we refuse is a dump of 4-grams from an unnamed event stream. If you cannot name the actor and the grain, stop mining.

Module 03 of Sequence Atlas is entirely about this restraint. Feature Gravity Lab skips mining altogether and starts from return rates — useful when your taxonomy is not yet ready for sequences.