📋 KEY INSIGHTS
- NLP applications sit on a spectrum from rule-based to fully neural. For many production problems — email routing, basic intent classification, keyword extraction — a TF-IDF vectoriser with a logistic regression classifier outperforms a fine-tuned BERT model in speed, cost, and maintainability.
- Sentiment analysis is deceptively difficult: aspect-level sentiment (positive about battery life, negative about camera) is substantially harder than document-level sentiment, and most off-the-shelf tools only solve the simpler version.
- Named Entity Recognition (NER) models are domain-sensitive — a model trained on news articles performs poorly on clinical notes or legal contracts. Domain-specific NER almost always requires fine-tuning on in-domain labelled data.
- Text classification at scale requires careful attention to class imbalance, label quality, and distribution shift — a model that achieves 92% accuracy on a balanced test set may perform at 60% on the live imbalanced production distribution.
- Zero-shot and few-shot classification using LLMs (asking GPT-4 or Claude to classify without training examples) is competitive with supervised models for many tasks and should be the first baseline attempted before investing in labelled data collection.
- The most impactful NLP applications in enterprise settings are document processing (contract analysis, invoice extraction, medical record summarisation) rather than the chatbot use cases that receive the most media attention.
Natural language processing has undergone two revolutions in the past decade. The first was word embeddings (word2vec, GloVe, FastText), which replaced one-hot encoding with dense semantic representations and made NLP models generalisable across domains for the first time. The second was the Transformer architecture and pre-trained language models (BERT, GPT, T5), which brought transfer learning to NLP — enabling fine-tuned models that outperform task-specific models trained from scratch on orders of magnitude more data. Today, data scientists working on text problems have access to tools that were science fiction ten years ago. This guide focuses on three of the most common and commercially valuable NLP applications — sentiment analysis, named entity recognition, and text classification — with a practical orientation toward production deployment.
Sentiment Analysis — Beyond Positive and Negative
Sentiment analysis (also called opinion mining) automatically identifies the emotional tone expressed in text. It is one of the oldest and most commercially deployed NLP tasks: companies analyse customer reviews, social media mentions, support tickets, and survey responses to understand how customers feel about their products, features, and service. Despite its apparent simplicity, production sentiment analysis is substantially harder than the academic benchmarks suggest.
The most common taxonomy distinguishes three levels of granularity. Document-level sentiment assigns a single polarity (positive, negative, neutral) or a numeric score to an entire document — a review, a tweet, a support ticket. This is the easiest level and is well-solved by off-the-shelf tools and fine-tuned BERT models. Sentence-level sentiment assigns polarity to individual sentences within a document — relevant when a document contains mixed sentiment. Aspect-level sentiment analysis (ABSA) identifies both the aspect being discussed (product, service, price, delivery) and the sentiment expressed toward each aspect — enabling insights like “customers are positive about product quality but strongly negative about shipping times.” ABSA is substantially harder and requires more sophisticated models or fine-tuning on domain-specific data.
The practical limitations of sentiment analysis are important to communicate to stakeholders. Sarcasm and irony (“Oh great, another update that breaks everything”) consistently fool sentiment models. Domain-specific language — medical, legal, financial — contains vocabulary that carries different connotations than in everyday text. Neutral statements that are factually negative (“The battery lasted 2 hours before needing a charge”) are often misclassified as neutral by models trained on opinion data. Multilingual sentiment requires either a multilingual model (XLM-RoBERTa) or separate models per language, not a single English-trained model applied to translated text.
| Approach | Best For | Accuracy (typical) | Training Data Required | Inference Speed |
|---|---|---|---|---|
| Lexicon-based (VADER, SentiWordNet) | Social media, quick baseline | 70–80% | None | Very fast (rule-based) |
| TF-IDF + Logistic Regression | Domain-specific with labelled data | 80–88% | 1,000–10,000 examples | Very fast |
| Fine-tuned BERT / RoBERTa | General or domain with fine-tuning | 90–95% | 500–5,000 examples | Medium (GPU recommended) |
| Zero-shot LLM (GPT-4, Claude) | Fast prototyping, rare categories | 85–92% | None | Slow (API latency) |
| Aspect-Based (ABSA model) | Product/service feedback analysis | 75–85% per aspect | 5,000–50,000 examples | Medium |
Named Entity Recognition — Extracting Structured Information from Text
Named Entity Recognition (NER) identifies and classifies named entities in text — people, organisations, locations, dates, monetary values, products, medical terms, and any custom entity type relevant to your domain. It is a foundational component in document processing pipelines, knowledge graph construction, information extraction, and search. A well-functioning NER model turns unstructured text into structured data that can be queried, aggregated, and analysed.
Standard NER models cover a small set of generic entity types: PERSON, ORG (organisation), LOC (location), DATE, TIME, MONEY, PERCENT, and GPE (geopolitical entity). These models — available in spaCy, NLTK, and Hugging Face — work well for news articles, general web text, and domains similar to their training data. They perform poorly in specialised domains. A clinical NER model needs to extract medication names, dosages, conditions, symptoms, and anatomical structures — entity types completely absent from general NER training data. A legal NER model must extract parties, jurisdictions, case numbers, and contractual clauses. Building domain NER requires labelling in-domain text with your custom entity types and fine-tuning a base model (BERT, RoBERTa) on those labels.
The standard evaluation metric for NER is the F1 score at the span level — an entity is counted as correct only if both its boundaries (start and end token positions) and its type are predicted correctly. This is stricter than token-level accuracy and better reflects real-world usefulness: a model that correctly identifies “Tata Consultancy” but misses “Services” produces a wrong entity boundary and is not useful for downstream processing, even though it got most tokens right.
| Domain | Key Entity Types | Recommended Base Model | Typical Labelled Data Needed |
|---|---|---|---|
| General (news, web) | PERSON, ORG, LOC, DATE, MONEY | spaCy en_core_web_trf or spaCy large | Use pre-trained (no fine-tuning) |
| Clinical / Medical | Disease, Drug, Dosage, Symptom, Anatomy | BioBERT or ClinicalBERT | 2,000–10,000 annotated sentences |
| Legal | Party, Jurisdiction, Case Number, Clause | Legal-BERT | 2,000–8,000 annotated sentences |
| Finance | Ticker, Company, Financial Metric, Date | FinBERT | 1,000–5,000 annotated sentences |
| E-commerce / Retail | Product, Brand, Colour, Size, Material | RoBERTa-base fine-tuned | 3,000–15,000 annotated examples |
| Custom domain | Domain-specific types | RoBERTa-base or domain BERT | 1,000–20,000 examples |
Text Classification — From Spam Filtering to Intent Detection
Text classification assigns one or more predefined categories to a text document. It is the most widely deployed NLP task: spam filtering, support ticket routing, news categorisation, content moderation, intent detection in chatbots, and document triage all reduce to text classification. The right modelling approach depends on the number of classes, class balance, labelled data availability, and latency requirements.
For binary classification with abundant labelled data and a standard domain (similar to news or social media), a fine-tuned BERT or DistilBERT model is the current best practice. DistilBERT achieves 97% of BERT’s performance at 60% of its size and 2× the inference speed — it is the standard choice for production deployments where latency matters. For multi-class classification with many classes (100+) and limited data per class, few-shot techniques — training a contrastive learning model (SetFit) on 8–32 examples per class — significantly outperform standard fine-tuning in low-data regimes.
Zero-shot classification deserves special attention. Using an LLM (via API) with a well-designed prompt, you can classify text into categories that were never seen during the model’s training. This approach requires no labelled data and no training — you simply prompt the model with the text and the categories, and it returns a prediction. For many real-world business problems with fewer than 20 categories and tolerance for API latency, zero-shot classification is the fastest path to a working prototype and is often competitive with supervised models when the category definitions are unambiguous. The limitation is cost (API calls for every prediction) and latency (~500ms+ per request), which makes it unsuitable for high-volume, latency-sensitive applications.
| Scenario | Recommended Approach | Why |
|---|---|---|
| Binary, 10k+ labelled examples | Fine-tuned DistilBERT | High accuracy, fast inference |
| Multi-class, 100+ classes, few examples | SetFit (few-shot contrastive) | Works well with 8–32 examples per class |
| New task, no labelled data | Zero-shot LLM classification | No annotation required, works immediately |
| Simple categories, high volume | TF-IDF + Logistic Regression | Fastest inference, lowest cost |
| Hierarchical taxonomy | Top-down classifier cascade | Coarse → fine prediction reduces errors |
| Multilingual | Fine-tuned XLM-RoBERTa | 100+ languages, single model |
| Noisy, short text (tweets, reviews) | FastText + cleaning pipeline | Robust to informal language at speed |
✦ SUMMARIZE THIS ARTICLE WITH AI
The full NLP preprocessing pipeline — tokenisation, stopword removal, stemming, lemmatisation, and word embeddings — that feeds into these application models is covered in our NLP Pipeline guide. The Transformer architecture and BERT pre-training that underlies modern NLP classification is explained in our Transformers guide. For common NLP interview questions covering BERT, attention, and sequence modelling, see our NLP Interview Q&A. Generative AI applications built on top of these NLP foundations are in our Generative AI guide.



