Thursday, October 1, 2026
HomeData ScienceGenerative AI for Data Scientists – LLMs, RAG, Fine-Tuning and Prompt Engineering

Generative AI for Data Scientists – LLMs, RAG, Fine-Tuning and Prompt Engineering

Table of Content

📋 KEY INSIGHTS

  • Generative AI refers to models that learn the distribution of training data and can generate new samples from it β€” including large language models (LLMs), diffusion models, and variational autoencoders.
  • LLMs like GPT-4, Claude, and Llama are trained on massive text corpora with next-token prediction, then aligned with human preferences via RLHF (Reinforcement Learning from Human Feedback) to follow instructions.
  • RAG (Retrieval-Augmented Generation) combines a retrieval system (vector database) with a generative model to ground responses in specific documents β€” solving the knowledge cutoff and hallucination problems of standalone LLMs.
  • Fine-tuning adapts a pre-trained LLM to a specific domain or task using a small dataset. Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA modify only a small fraction of parameters, reducing GPU memory by 90%+.
  • Prompt engineering β€” crafting the input to elicit the desired output β€” is a high-leverage skill that requires no model training. Techniques like chain-of-thought, few-shot examples, and system prompts can dramatically improve output quality.
  • Evaluating generative AI outputs is fundamentally harder than evaluating discriminative models β€” there is no single correct answer. LLM-as-judge, human evaluation, and task-specific metrics (BLEU, ROUGE, BERTScore) each capture different dimensions of quality.

Generative AI has shifted from a research curiosity to a production staple in the span of three years. Every data science team is now expected to understand how large language models work, when to use RAG versus fine-tuning, how to evaluate outputs, and how to build applications on top of these systems. This is not a small supplement to the traditional ML skill set β€” for many teams, generative AI has become the primary focus. This guide gives you the conceptual and practical foundations: how LLMs are built, the anatomy of a RAG pipeline, the trade-offs between prompting, RAG, and fine-tuning, and how to evaluate systems that generate rather than classify.

How Large Language Models Work

A large language model is, at its core, a very large neural network β€” specifically a Transformer decoder β€” trained to predict the next token in a sequence. Given the context “The capital of France is”, the model assigns probabilities to every token in its vocabulary and samples “Paris” (or another plausible continuation). That single training objective β€” next-token prediction β€” turns out to be sufficient to learn grammar, world knowledge, reasoning, and even code, when applied at sufficient scale (billions of parameters, trillions of training tokens). This is not magic; it is compression: to predict next tokens accurately, the model must develop internal representations that capture the structure of language and knowledge.

The training pipeline for a modern instruction-following LLM has three stages. Pre-training trains on a massive corpus of internet text (Common Crawl, books, code, Wikipedia) with next-token prediction on trillions of tokens. This is by far the most compute-intensive step β€” GPT-3 required approximately 3,640 petaflop-days of compute. Supervised Fine-Tuning (SFT) trains on a smaller dataset of high-quality (prompt, desired response) pairs created by human labellers, teaching the model to follow instructions rather than just continue text. Reinforcement Learning from Human Feedback (RLHF) further aligns the model with human preferences: human raters rank model outputs, a reward model is trained on these rankings, and the LLM is updated via PPO (Proximal Policy Optimisation) to generate outputs that score highly on the reward model. The result is a model that is helpful, harmless, and honest β€” the three H’s that define alignment.

Model FamilyArchitectureTraining DataContext WindowOpen Source?Best For
GPT-4 / GPT-4oTransformer decoder (MoE)Proprietary128k tokensNo (API only)General reasoning, multimodal
Claude 3.5 / SonnetTransformer decoderProprietary200k tokensNo (API only)Long context, coding, analysis
Llama 3.1 / 3.3Transformer decoderPublic + proprietary128k tokensYes (weights)Self-hosted, fine-tuning
Mistral / MixtralTransformer decoder (MoE)Proprietary32k tokensYes (weights)Efficient, fast inference
Gemini 1.5 ProMultimodal TransformerProprietary1M tokensNo (API only)Very long documents, video
Stable Diffusion 3Diffusion + TransformerLAIONN/A (image)YesImage generation

RAG β€” Retrieval-Augmented Generation

Abstract 3D composition with red shapes and rockwell automation logo
Photo by Brecht Corbeel on Unsplash

The two fundamental limitations of standalone LLMs are a knowledge cutoff (the model does not know about events after its training data was collected) and hallucination (the model confabulates plausible-sounding but incorrect facts, especially for specific or obscure topics). RAG addresses both by giving the model access to a retrieval system at inference time.

In a RAG pipeline, documents are first chunked into passages and embedded into dense vectors using an embedding model (OpenAI text-embedding-3-large, or an open-source model like BGE or E5). These vectors are stored in a vector database (Pinecone, Weaviate, Qdrant, or pgvector for PostgreSQL). At query time, the user’s question is embedded with the same model, and the nearest-neighbour passages are retrieved. The retrieved passages are prepended to the LLM’s prompt as context, and the model generates a response grounded in those passages. The model is instructed to cite its sources and to say “I don’t know” if the answer is not in the retrieved context.

RAG ComponentWhat It DoesCommon ToolsKey Design Decision
Document LoaderIngest PDFs, web pages, databasesLangChain, LlamaIndex, UnstructuredWhich sources to index
ChunkerSplit documents into retrievable passagesRecursiveTextSplitter, semantic splitterChunk size and overlap (512 tokens, 50 overlap typical)
Embedding ModelConvert text to dense vectorstext-embedding-3-large, BGE-M3, E5Dimension, multilingual support
Vector StoreStore and search embeddingsPinecone, Qdrant, Weaviate, pgvectorScale, latency, managed vs self-hosted
RetrieverFind top-k relevant chunksANN search (HNSW), BM25, hybridk value; dense vs sparse vs hybrid retrieval
LLM GeneratorSynthesise answer from contextGPT-4o, Claude, Llama, MistralSystem prompt, citation format, temperature
EvaluatorMeasure faithfulness and relevanceRAGAS, TruLens, DeepEvalFaithfulness, answer relevance, context recall

The most important RAG design decisions are chunking strategy, retrieval method, and reranking. Fixed-size chunking is simple but splits semantic units; sentence or paragraph chunking preserves meaning better. Hybrid retrieval β€” combining dense embedding similarity with sparse BM25 keyword matching β€” significantly outperforms either alone, especially for queries with specific technical terms or names. A cross-encoder reranker (a smaller model that scores each retrieved chunk against the query) can be applied after initial retrieval to reorder the top-k passages before passing them to the generator, improving precision substantially.

Prompting vs RAG vs Fine-Tuning β€” When to Use Each

Data scientists frequently ask which approach to use when adapting an LLM to a new task. The answer depends on what the gap is between the base model’s behaviour and the desired behaviour. These three approaches address different gaps and have very different cost, complexity, and maintenance profiles.

Prompt engineering is the fastest and cheapest approach. It changes the model’s input without changing the model’s weights. Effective prompting techniques include: providing a clear system prompt that defines the model’s role and constraints; giving 2–5 few-shot examples of the desired input/output format; using chain-of-thought prompting (“think step by step”) for reasoning tasks; and breaking complex tasks into a chain of simpler prompts (prompt chaining). Prompt engineering should always be tried first. It costs nothing except experimentation time and is fully reversible. If a well-crafted prompt achieves the desired performance, there is no justification for the complexity and cost of fine-tuning.

RAG is the right choice when the model needs access to information that was not in its training data (proprietary documents, recent events, internal knowledge bases) or when responses must be traceable to specific source documents (regulated industries, customer support with cited policies). RAG does not change the model’s behaviour or reasoning β€” it only augments its context window with retrieved information. The failure modes of RAG are retrieval failures (the relevant chunk is not retrieved) rather than model failures.

Fine-tuning is the right choice when you need the model to consistently adopt a specific format, style, or domain vocabulary that cannot be reliably achieved through prompting; when you need to specialise the model for a narrow task where it can outperform a larger general model; or when inference latency and cost require a smaller, faster model. Fine-tuning with LoRA (Low-Rank Adaptation) adds small trainable rank-decomposition matrices to the attention layers, modifying only ~0.1% of parameters while achieving 90%+ of the performance of full fine-tuning. This makes fine-tuning feasible on a single consumer GPU for models up to 13B parameters.

ApproachChanges Model Weights?Needs Labelled Data?CostBest For
Prompt EngineeringNoNoNear zeroFormat, tone, reasoning style changes
RAGNoNo (for setup)Low–Medium (infra)Private / recent knowledge, citation
Fine-Tuning (LoRA)Yes (partially)Yes (100–10k pairs)Medium (GPU hours)Domain style, narrow task specialisation
Full Fine-TuningYes (all weights)Yes (10k+ pairs)High (multi-GPU days)Significant capability changes, alignment
Pre-training from scratchYes (all weights)Yes (billions of tokens)Very High (millions USD)Domain-specific LLM (BioBERT, CodeLlama)

✦ SUMMARIZE THIS ARTICLE WITH AI

The Transformer architecture underlying all LLMs β€” self-attention, multi-head attention, positional encoding β€” is explained from scratch in our Transformers and Attention guide. Fine-tuning concepts including transfer learning and domain adaptation are in our Transfer Learning guide. NLP preprocessing pipelines that feed into LLM applications are covered in our NLP Pipeline guide. For the full NLP interview Q&A including BERT, GPT, and attention mechanism questions, see our NLP Interview Q&A.

Leave feedback about this

  • Rating

Durgesh Kekare
Durgesh Kekarehttps://www.dataexpertise.in
Durgesh Kekare is a data science educator and founder of DataExpertise.in. With expertise in Python, machine learning, and analytics, he helps 10,000+ learners break into data careers.

Latest Posts

List of Categories