1038 - Natural Language Processing 英授 Taught in English

Natural Language Processing

教育目標 Course Target

This course bridges statistical NLP, deep neural sequence modeling, and modern Large Language Models (LLMs). The course equips students with mathematical theory and hands-on PyTorch skills to build, evaluate, and deploy production NLP systems from first principles.
Learning Outcomes:
1. Foundations: Analyze discrete text signals, linguistic ambiguity (6 tiers), and statistical power laws (Zipf, Heaps) to build robust Unicode NFKC pipelines.
2. Algorithmic Mastery: Implement core algorithms from scratch: BPE subword tokenizers, smoothed N-gram LMs, dynamic programming Viterbi POS taggers, and Shift-Reduce dependency parsers.
3. Neural Sequence Models: Build and train deep neural networks in PyTorch, including Bengio (2003) neural LMs, custom 4-gate LSTM cells, and Seq2Seq with Bahdanau attention.
4. Transformers & LLMs: Implement Multi-Head Attention from raw tensor math, fine-tune pre-trained BERT models, build Vector RAG pipelines, and execute LoRA fine-tuning.
5. Systems & Safety: Benchmark models (Perplexity, F1, BLEU), audit demographic fairness parity, and deploy multi-tier AI safety guardrails and INT8 quantization.

This course bridges statistical NLP, deep neural sequence modeling, and modern Large Language Models (LLMs). The course equips students with mathematical theory and hands-on PyTorch skills to build, evaluate, and deploy production NLP systems from first principles.
Learning Outcomes:
1. Foundations: Analyze discrete text signals, linguistic ambiguity (6 tiers), and statistical power laws (Zipf, Heaps) to build robust Unicode NFKC pipelines.
2. Algorithmic Mastery: Implement core algorithms from scratch: BPE subword tokenizers, smoothed N-gram LMs, dynamic programming Viterbi POS taggers, and Shift-Reduce dependency parsers.
3. Neural Sequence Models: Build and train deep neural networks in PyTorch, including Bengio (2003) neural LMs, custom 4-gate LSTM cells, and Seq2Seq with Bahdanau attention.
4. Transformers & LLMs: Implement Multi-Head Attention from raw tensor math, fine-tune pre-trained BERT models, build Vector RAG pipelines, and execute LoRA fine-tuning.
5. Systems & Safety: Benchmark models (Perplexity, F1, BLEU), audit demographic fairness parity, and deploy multi-tier AI safety guardrails and INT8 quantization.

參考書目 Reference Books

1. Primary Textbook (指定教科書 / Required Textbook)

a. Kristiani, Endah. (2026). Applied Natural Language Processing: A Hands-on Engineering Approach with PyTorch and Transformers (Course Notes & Lab Manual). Tunghai University. (Provided in PDF).

b. Jurafsky, Daniel, & Martin, James H. (2024). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition (3rd ed. draft). Stanford University / Pearson.
Online Free Access: https://web.stanford.edu/~jurafsky/slp3/

2. Recommended Reference Books (參考書籍 / Reference Materials)

a. Tunstall, Lewis, von Werra, Leandro, & Wolf, Thomas. (2022).
Natural Language Processing with Transformers: Building Language Applications with Hugging Face (Revised 2nd ed.). O'Reilly Media. ISBN: 978-1098136796.

b. Eisenstein, Jacob. (2019).
Introduction to Natural Language Processing. MIT Press. ISBN: 978-0262042840.

c. Goldberg, Yoav. (2017).
Neural Network Methods in Natural Language Processing (Synthesis Lectures on Human Language Technologies). Morgan & Claypool Publishers. ISBN: 978-1627052986.

d. Vaswani, Ashish, et al. (2017).
"Attention Is All You Need." Advances in Neural Information Processing Systems (NeurIPS 2017).

e. Hu, Edward J., et al. (2021).
"LoRA: Low-Rank Adaptation of Large Language Models." arXiv preprint arXiv:2106.09685.

f. Lewis, Patrick, et al. (2020).
"Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS 2020.

3. Online Resources & Documentation (線上學習資源與工具文件)

a. PyTorch Official Tutorials & Documentation: https://pytorch.org/tutorials/

b. Hugging Face Transformers & Datasets Documentation: https://huggingface.co/docs

c. LangChain & LlamaIndex Official Documentation: https://docs.langchain.com / d.https://docs.llamaindex.ai

e. Stanford CS224N: Natural Language Processing with Deep Learning: https://web.stanford.edu/class/cs224n/

1. Primary Textbook (Specified Textbook / Required Textbook)

a. Kristiani, Endah. (2026). Applied Natural Language Processing: A Hands-on Engineering Approach with PyTorch and Transformers (Course Notes & Lab Manual). Tunghai University. (Provided in PDF).

b. Jurafsky, Daniel, & Martin, James H. (2024). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition (3rd ed. draft). Stanford University / Pearson.
Online Free Access: https://web.stanford.edu/~jurafsky/slp3/

2. Recommended Reference Books (Reference Materials)

a. Tunstall, Lewis, von Werra, Leandro, & Wolf, Thomas. (2022).
Natural Language Processing with Transformers: Building Language Applications with Hugging Face (Revised 2nd ed.). O'Reilly Media. ISBN: 978-1098136796.

b. Eisenstein, Jacob. (2019).
Introduction to Natural Language Processing. MIT Press. ISBN: 978-0262042840.

c. Goldberg, Yoav. (2017).
Neural Network Methods in Natural Language Processing (Synthesis Lectures on Human Language Technologies). Morgan & Claypool Publishers. ISBN: 978-1627052986.

d. Vaswani, Ashish, et al. (2017).
"Attention Is All You Need." Advances in Neural Information Processing Systems (NeurIPS 2017).

e. Hu, Edward J., et al. (2021).
"LoRA: Low-Rank Adaptation of Large Language Models." arXiv preprint arXiv:2106.09685.

f. Lewis, Patrick, et al. (2020).
"Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS 2020.

3. Online Resources & Documentation (online learning resources and tool files)

a. PyTorch Official Tutorials & Documentation: https://pytorch.org/tutorials/

b. Hugging Face Transformers & Datasets Documentation: https://huggingface.co/docs

c. LangChain & LlamaIndex Official Documentation: https://docs.langchain.com / d.https://docs.llamaindex.ai

e. Stanford CS224N: Natural Language Processing with Deep Learning: https://web.stanford.edu/class/cs224n/

評分方式 Grading

評分項目
Grading Method
配分比例
Percentage
說明
Description
Hands-On Lab Assignments
Hands-On Lab Assignments
20 Weekly PyTorch and Python coding labs evaluating algorithmic implementation from scratch (BPE tokenization, N-grams, 4-gate LSTM, and Attention) alongside automated unit tests and code quality.
Midterm Project Milestone & Exam
Midterm Project Milestone & Exam
30 Assesses theoretical understanding and practical programming proficiency across classical statistical NLP, word vector embeddings, language modeling, and neural sequence architectures.
Final Capstone Project & Presentation
Final Capstone Project & Presentation
40 Team-based end-to-end NLP/LLM engineering project involving problem formulation, RAG/LoRA model implementation, quantitative performance evaluation, bias/safety audit, technical report, and live demo.
Class Participation & Socratic Discussions
Class Participation & Socratic Discussions
10 Evaluates active engagement during in-class interactive code walkthroughs, Socratic design discussions, architectural critique sessions, and peer feedback.

授課大綱 Course Plan

點擊下方連結查看詳細授課大綱
Click the link below to view the detailed course plan

查看授課大綱 View Course Plan

相似課程 Related Courses

無相似課程 No related courses found

課程資訊 Course Information

基本資料 Basic Information

  • 課程代碼 Course Code: 1038
  • 學分 Credit: 3-0
  • 上課時間 Course Time:
    Thursday/6,7,8[C202]
  • 授課教師 Teacher:
    恩達
  • 修課班級 Class:
    資工系2-4
  • 選課備註 Memo:
    資工系國際組;全英授課
選課狀態 Enrollment Status

目前選課人數 Current Enrollment: 1 人

交換生/外籍生選課登記

請點選上方按鈕加入登記清單,再等候任課教師審核。
Add this class to your wishlist by clicking the button above.