Get in Touch
 Duration 21 hours

Course Outline

Comprehensive Training Agenda

  1. Introduction to NLP
    • Foundations of NLP
    • Overview of NLP Frameworks
    • Commercial applications of NLP
    • Web data scraping techniques
    • Utilizing various APIs to fetch text data
    • Managing text corpora: storage of content and relevant metadata
    • Benefits of Python and an introductory NLTK session
  2. Practical Insights into Corpora and Datasets
    • The necessity of a corpus
    • Corpus analysis methods
    • Categories of data attributes
    • File formats for corpora
    • Preparing datasets for NLP applications
  3. Sentence Structure Analysis
    • Core NLP components
    • Natural language understanding
    • Morphological analysis: stems, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Managing ambiguity
  4. Text Data Preprocessing
    • Corpus: Raw text processing
      • Sentence tokenization
      • Stemming of raw text
      • Lemmatization of raw text
      • Removal of stop words
    • Corpus: Raw sentences
      • Word tokenization
      • Word lemmatization
    • Constructing Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Customized preprocessing strategies
  5. Text Data Analysis
    • Fundamental NLP features
      • Parsers and parsing techniques
      • Part-of-speech (POS) tagging and taggers
      • Named entity recognition
      • N-grams
      • Bag of words model
    • Statistical aspects of NLP
      • Linear algebra concepts for NLP
      • Probabilistic theory in NLP
      • TF-IDF
      • Vectorization
      • Encoders and Decoders
      • Normalization
      • Probabilistic Models
    • Advanced feature engineering in NLP
      • Introduction to word2vec
      • Components of the word2vec model
      • Underlying logic of the word2vec model
      • Extensions of the word2vec concept
      • Practical applications of the word2vec model
    • Case study: Implementing automatic text summarization using the Bag of Words approach with simplified and exact Luhn's algorithms
  6. Document Clustering, Classification, and Topic Modelling
    • Document clustering and pattern mining (including hierarchical clustering and k-means)
    • Document comparison and classification using TFIDF, Jaccard, and cosine distance metrics
    • Document classification using Naïve Bayes and Maximum Entropy
  7. Identifying Key Text Elements
    • Dimensionality reduction: Principal Component Analysis, Singular Value Decomposition, and Non-negative Matrix Factorization
    • Topic modelling and information retrieval via Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis, and Advanced Topic Modelling
    • Positive vs. negative: Measuring sentiment intensity
    • Item Response Theory
    • Part-of-speech tagging applications: Identifying people, places, and organizations in text
    • Advanced topic modelling: Latent Dirichlet Allocation
  9. Case Studies
    • Analyzing unstructured user reviews
    • Sentiment classification and visualization of product review data
    • Mining search logs to identify usage patterns
    • Text classification
    • Topic modelling

Requirements

Familiarity with NLP fundamentals and an understanding of AI applications in business contexts

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories