1038 - Natural Language Processing 英授 Taught in English
Natural Language Processing
教育目標 Course Target
This course bridges statistical NLP, deep neural sequence modeling, and modern Large Language Models (LLMs). The course equips students with mathematical theory and hands-on PyTorch skills to build, evaluate, and deploy production NLP systems from first principles.
Learning Outcomes:
1. Foundations: Analyze discrete text signals, linguistic ambiguity (6 tiers), and statistical power laws (Zipf, Heaps) to build robust Unicode NFKC pipelines.
2. Algorithmic Mastery: Implement core algorithms from scratch: BPE subword tokenizers, smoothed N-gram LMs, dynamic programming Viterbi POS taggers, and Shift-Reduce dependency parsers.
3. Neural Sequence Models: Build and train deep neural networks in PyTorch, including Bengio (2003) neural LMs, custom 4-gate LSTM cells, and Seq2Seq with Bahdanau attention.
4. Transformers & LLMs: Implement Multi-Head Attention from raw tensor math, fine-tune pre-trained BERT models, build Vector RAG pipelines, and execute LoRA fine-tuning.
5. Systems & Safety: Benchmark models (Perplexity, F1, BLEU), audit demographic fairness parity, and deploy multi-tier AI safety guardrails and INT8 quantization.
This course bridges statistical NLP, deep neural sequence modeling, and modern Large Language Models (LLMs). The course equips students with mathematical theory and hands-on PyTorch skills to build, evaluate, and deploy production NLP systems from first principles.
Learning Outcomes:
1. Foundations: Analyze discrete text signals, linguistic ambiguity (6 tiers), and statistical power laws (Zipf, Heaps) to build robust Unicode NFKC pipelines.
2. Algorithmic Mastery: Implement core algorithms from scratch: BPE subword tokenizers, smoothed N-gram LMs, dynamic programming Viterbi POS taggers, and Shift-Reduce dependency parsers.
3. Neural Sequence Models: Build and train deep neural networks in PyTorch, including Bengio (2003) neural LMs, custom 4-gate LSTM cells, and Seq2Seq with Bahdanau attention.
4. Transformers & LLMs: Implement Multi-Head Attention from raw tensor math, fine-tune pre-trained BERT models, build Vector RAG pipelines, and execute LoRA fine-tuning.
5. Systems & Safety: Benchmark models (Perplexity, F1, BLEU), audit demographic fairness parity, and deploy multi-tier AI safety guardrails and INT8 quantization.
參考書目 Reference Books
1. Primary Textbook (指定教科書 / Required Textbook)
a. Kristiani, Endah. (2026). Applied Natural Language Processing: A Hands-on Engineering Approach with PyTorch and Transformers (Course Notes & Lab Manual). Tunghai University. (Provided in PDF).
b. Jurafsky, Daniel, & Martin, James H. (2024). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition (3rd ed. draft). Stanford University / Pearson.
Online Free Access: https://web.stanford.edu/~jurafsky/slp3/
2. Recommended Reference Books (參考書籍 / Reference Materials)
a. Tunstall, Lewis, von Werra, Leandro, & Wolf, Thomas. (2022).
Natural Language Processing with Transformers: Building Language Applications with Hugging Face (Revised 2nd ed.). O'Reilly Media. ISBN: 978-1098136796.
b. Eisenstein, Jacob. (2019).
Introduction to Natural Language Processing. MIT Press. ISBN: 978-0262042840.
c. Goldberg, Yoav. (2017).
Neural Network Methods in Natural Language Processing (Synthesis Lectures on Human Language Technologies). Morgan & Claypool Publishers. ISBN: 978-1627052986.
d. Vaswani, Ashish, et al. (2017).
"Attention Is All You Need." Advances in Neural Information Processing Systems (NeurIPS 2017).
e. Hu, Edward J., et al. (2021).
"LoRA: Low-Rank Adaptation of Large Language Models." arXiv preprint arXiv:2106.09685.
f. Lewis, Patrick, et al. (2020).
"Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS 2020.
3. Online Resources & Documentation (線上學習資源與工具文件)
a. PyTorch Official Tutorials & Documentation: https://pytorch.org/tutorials/
b. Hugging Face Transformers & Datasets Documentation: https://huggingface.co/docs
c. LangChain & LlamaIndex Official Documentation: https://docs.langchain.com / d.https://docs.llamaindex.ai
e. Stanford CS224N: Natural Language Processing with Deep Learning: https://web.stanford.edu/class/cs224n/
1. Primary Textbook (Specified Textbook / Required Textbook)
a. Kristiani, Endah. (2026). Applied Natural Language Processing: A Hands-on Engineering Approach with PyTorch and Transformers (Course Notes & Lab Manual). Tunghai University. (Provided in PDF).
b. Jurafsky, Daniel, & Martin, James H. (2024). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition (3rd ed. draft). Stanford University / Pearson.
Online Free Access: https://web.stanford.edu/~jurafsky/slp3/
2. Recommended Reference Books (Reference Materials)
a. Tunstall, Lewis, von Werra, Leandro, & Wolf, Thomas. (2022).
Natural Language Processing with Transformers: Building Language Applications with Hugging Face (Revised 2nd ed.). O'Reilly Media. ISBN: 978-1098136796.
b. Eisenstein, Jacob. (2019).
Introduction to Natural Language Processing. MIT Press. ISBN: 978-0262042840.
c. Goldberg, Yoav. (2017).
Neural Network Methods in Natural Language Processing (Synthesis Lectures on Human Language Technologies). Morgan & Claypool Publishers. ISBN: 978-1627052986.
d. Vaswani, Ashish, et al. (2017).
"Attention Is All You Need." Advances in Neural Information Processing Systems (NeurIPS 2017).
e. Hu, Edward J., et al. (2021).
"LoRA: Low-Rank Adaptation of Large Language Models." arXiv preprint arXiv:2106.09685.
f. Lewis, Patrick, et al. (2020).
"Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS 2020.
3. Online Resources & Documentation (online learning resources and tool files)
a. PyTorch Official Tutorials & Documentation: https://pytorch.org/tutorials/
b. Hugging Face Transformers & Datasets Documentation: https://huggingface.co/docs
c. LangChain & LlamaIndex Official Documentation: https://docs.langchain.com / d.https://docs.llamaindex.ai
e. Stanford CS224N: Natural Language Processing with Deep Learning: https://web.stanford.edu/class/cs224n/
評分方式 Grading
| 評分項目 Grading Method |
配分比例 Percentage |
說明 Description |
|---|---|---|
|
Hands-On Lab Assignments Hands-On Lab Assignments |
20 | Weekly PyTorch and Python coding labs evaluating algorithmic implementation from scratch (BPE tokenization, N-grams, 4-gate LSTM, and Attention) alongside automated unit tests and code quality. |
|
Midterm Project Milestone & Exam Midterm Project Milestone & Exam |
30 | Assesses theoretical understanding and practical programming proficiency across classical statistical NLP, word vector embeddings, language modeling, and neural sequence architectures. |
|
Final Capstone Project & Presentation Final Capstone Project & Presentation |
40 | Team-based end-to-end NLP/LLM engineering project involving problem formulation, RAG/LoRA model implementation, quantitative performance evaluation, bias/safety audit, technical report, and live demo. |
|
Class Participation & Socratic Discussions Class Participation & Socratic Discussions |
10 | Evaluates active engagement during in-class interactive code walkthroughs, Socratic design discussions, architectural critique sessions, and peer feedback. |
授課大綱 Course Plan
點擊下方連結查看詳細授課大綱
Click the link below to view the detailed course plan
相似課程 Related Courses
無相似課程 No related courses found
課程資訊 Course Information
基本資料 Basic Information
- 課程代碼 Course Code: 1038
- 學分 Credit: 3-0
-
上課時間 Course Time:Thursday/6,7,8[C202]
-
授課教師 Teacher:恩達
-
修課班級 Class:資工系2-4
-
選課備註 Memo:資工系國際組;全英授課
交換生/外籍生選課登記
請點選上方按鈕加入登記清單,再等候任課教師審核。
Add this class to your wishlist by clicking the button above.