Article
How to cluster open-text employee feedback into structural patterns
By Dr. Tim Hough LinkedIn
Founder, Hough and Associates, Inc.
Doctoral researcher of workplace frustration and engagement; author of The Frustration Condition (First Edition, 2026) and the 331-participant quantitative study of effort, frustration, and structural disengagement that grounds the framework.
Published · Last updated · 8 min read
To cluster open-text employee feedback in a way leadership can act on, route every statement into a fixed structural vocabulary — not a list of themes invented from this quarter's data — and force a recorded decision against each cluster. The fixed vocabulary is what gives you stability across cycles; the recorded decision is what closes the loop the survey leaves open.
Free-form thematic analysis is fine for a one-off research study. It is the wrong primitive for a recurring engagement program. The themes drift across analysts, across quarters, and across business units, and the drift erodes the comparability that makes the analysis useful in the first place.
Why the vocabulary must be fixed
If your cluster names change every quarter, you cannot tell whether a cluster grew, shrank, or was redistributed across new categories. Worse, you cannot tell leadership whether last quarter's decision against a cluster has stuck — because the cluster has been renamed.
The Frustration Condition framework fixes the vocabulary at five Architectures: Decision Bottlenecks, Approval Loops, Priority Churn, Role Ambiguity, and Unspoken Constraints. Every open-text statement is routed to exactly one of these. The framework is small enough to memorise and broad enough to absorb almost every structural friction a knowledge-work team will raise.
The clustering method, end-to-end
- Capture open-text via the Frustration Question — one prompt, no leading framing, no sentiment scale.
- Strip identifying details so the cluster reads against the team, not the person.
- Run the AI clustering pass — the model classifies each statement into one of the five Architectures, with a confidence score.
- Facilitator review — every cluster is editable, splittable, and re-mergeable. The facilitator is the decision-maker; the model is a fast first pass.
- Force a Three Doors decision per cluster — Remove, Defer With Clarity, or Accept — within 30 days.
- Report cluster size, recency, and decision status back to the team in language the team will recognise.
What the model is actually doing
The clustering model is not generating new categories. It is performing a constrained classification: for each open-text statement, which of five fixed labels best fits, with what confidence. That constraint is what keeps the output stable enough to compare across quarters, and what makes the model auditable — a misclassification is a discrete error against a published rubric, not a matter of taste.
Confidence scores below a published threshold are surfaced for facilitator review rather than auto-assigned. In our deployments, roughly 8–12% of statements fall into that band, which is a manageable review workload for a single-team Listen.
Common failure modes when clustering open-text
- Letting the model invent new themes per quarter — destroys cross-quarter comparability.
- Treating cluster size as the headline metric — a small cluster of severe statements often matters more than a large cluster of minor ones; use cluster size, recency, and decision status together.
- Closing the loop with a survey response rather than a recorded decision — restarts the same pattern next quarter.
- Hiding the open-text from leadership and reporting only the cluster names — leadership needs to read at least a sample to feel the structural shape.
- Reporting clusters without naming the door — the cluster is a question; the door is the answer.
Outside references on text clustering for HR data
Two useful outside references on the methodology side: Krippendorff's Content Analysis (the canonical text on coded thematic analysis at scale) and the Stanford NLP group's tutorials on constrained classification, which describe exactly the regime the AI cluster runs under here.
Covered in the book
The full treatment of this topic lives in Why Your Best People Stop Trying by Dr. Tim Hough.
Frequently asked
Common questions about How to cluster open-text employee feedback into structural patterns.
- Can we use our own custom themes instead of the five Architectures?
- You can, but you lose cross-quarter comparability and you lose the published vocabulary the line managers in your organisation will hear elsewhere. The five Architectures were chosen to be small enough to memorise and broad enough to absorb almost every structural friction a knowledge-work team raises.
- How accurate is the AI cluster?
- In production we see ~88–92% agreement with facilitator-reviewed cluster assignment, with the disagreement concentrated in the 8–12% of statements the model itself flags as low-confidence. Those go to the facilitator for review before any decision is recorded.
- What if a statement spans two Architectures?
- The facilitator can split the statement across two clusters at review time. In practice this is rare; most multi-Architecture statements have a primary structural cause, and forcing a primary classification is what makes the decision discipline workable.
- How often should we re-cluster?
- Cluster fresh after every Cross-Functional Listen — typically quarterly. Re-clustering historical data is rarely useful; the value is in the new statements and the decision-status check on the prior clusters.
Up & sideways