Streamline KD & QAD transformers Trainers #708

AAnoosheh · 2025-12-18T16:11:19Z

What does this PR do?

Type of change: ? Refactor and stabilization

Overview:

Enforce use of FSDP-2 on KD and QAD trainers in HF plugins/examples so that we can remove multiple restrictions

Usage

# Add a code snippet demonstrating how to use this

Testing

Before your PR is "Ready for review"

Make sure you read and follow Contributor guidelines and your commits are signed.
Is this change backward compatible?: Yes/No
Did you write any new necessary tests?: Yes/No
Did you add or update any necessary documentation?: Yes/No
Did you update Changelog?: Yes/No

Additional Information

copy-pr-bot · 2025-12-18T16:11:23Z

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

codecov · 2025-12-18T16:21:42Z

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 74.62%. Comparing base (03dc386) to head (7079dee).
⚠️ Report is 17 commits behind head on main.

Additional details and impacted files

@@            Coverage Diff             @@
##             main     #708      +/-   ##
==========================================
- Coverage   74.69%   74.62%   -0.07%     
==========================================
  Files         192      192              
  Lines       18946    18989      +43     
==========================================
+ Hits        14152    14171      +19     
- Misses       4794     4818      +24

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:

❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

kevalmorabia97 · 2025-12-19T14:57:37Z

examples/llm_qat/README.md

 > **_NOTE:_** `launch.sh` defaults to use `LlamaDecoderLayer` as the transformer layer class. If your model uses a different class, you need to pass `--fsdp_transformer_layer_cls_to_wrap <your_layer_class>` to the `launch.sh` script. For example, for `Qwen/Qwen3-8B`, specify `--fsdp_transformer_layer_cls_to_wrap Qwen3DecoderLayer` as an additional argument.

-> **_NOTE:_** The script defaults to using FSDP1. To use FSDP2, pass "--use_fsdp2 True" to the `launch.sh` script. Note that FSDP2 is less stable than FSDP1 currently. Use it with caution.
+> **_NOTE:_** The script defaults to using FSDP1. To use FSDP2, pass "--backend=fsdp2" to the `launch.sh` script. Note that FSDP2 is less stable than FSDP1 currently. Use it with caution.


Is this statement still valid? Note that FSDP2 is less stable than FSDP1 currently. Use it with caution.

I doubt it, but I don't have proof. I don't have proof that it is less stable either.

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

modelopt/torch/distill/plugins/huggingface.py

examples/llm_qat/README.md

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

examples/llm_qat/main.py

examples/llm_qat/README.md

modelopt/torch/quantization/plugins/transformers_trainer.py

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

modelopt/torch/quantization/plugins/transformers_trainer.py

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

modelopt/torch/quantization/plugins/transformers_trainer.py

realAsma · 2026-01-09T22:23:08Z

modelopt/torch/quantization/plugins/transformers_trainer.py

-        # Note: QAD doesn't work with FSDP wrapped model. We quantize model before the wrapper.
-        # The drawback is that we can't train a model that is bigger than a single GPU memory.
-        # And memory efficient loading doesn't work.
+        # Note: FSDP memory efficient loading doesn't work.


wont be needed if we do https://github.com/NVIDIA/Model-Optimizer/pull/708/files#r2677776721

Suggested change

# Note: FSDP memory efficient loading doesn't work.

modelopt/torch/quantization/plugins/transformers_trainer.py

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

realAsma

Looks great!!

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

AAnoosheh self-assigned this Dec 18, 2025

AAnoosheh marked this pull request as ready for review December 19, 2025 14:07

AAnoosheh requested review from a team as code owners December 19, 2025 14:07

AAnoosheh requested review from ChenhanYu, kevalmorabia97, kinjalpatel27, mxinO and realAsma December 19, 2025 14:07

kevalmorabia97 reviewed Dec 19, 2025

View reviewed changes

kevalmorabia97 approved these changes Dec 19, 2025

View reviewed changes

AAnoosheh added 6 commits December 22, 2025 06:42

Streamline KDTrainer for FSDP2

9e47b97

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

Refactor QADTrainer too

f2d97a1

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

No need for teacher_factory

62d5241

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

Enforce FSDP2 for KD/QAD

8b83f05

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

README and extra changes

65727c7

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

Remove outdated FSDP2 notes

190e4d2

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

AAnoosheh force-pushed the aanoosheh/kd-trainer-streamline branch from bde5788 to 190e4d2 Compare December 22, 2025 14:42

realAsma reviewed Dec 24, 2025

View reviewed changes

modelopt/torch/distill/plugins/huggingface.py Show resolved Hide resolved

realAsma reviewed Dec 24, 2025

View reviewed changes

examples/llm_qat/README.md Outdated Show resolved Hide resolved

AAnoosheh force-pushed the aanoosheh/kd-trainer-streamline branch from c4c0d19 to 9f3b0f8 Compare January 5, 2026 13:50

Update examples/llm_qat/README.md

12f1bfb

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

AAnoosheh force-pushed the aanoosheh/kd-trainer-streamline branch from 9f3b0f8 to 50cb0f4 Compare January 5, 2026 13:55

AAnoosheh added 2 commits January 6, 2026 06:55

Hide teacher during quantization

fc358d9

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

Allow other modes after kd_loss

796b023

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

AAnoosheh force-pushed the aanoosheh/kd-trainer-streamline branch from f6e1196 to 796b023 Compare January 6, 2026 14:56

realAsma reviewed Jan 7, 2026

View reviewed changes

examples/llm_qat/main.py Show resolved Hide resolved

realAsma reviewed Jan 7, 2026

View reviewed changes

examples/llm_qat/README.md Outdated Show resolved Hide resolved

realAsma reviewed Jan 7, 2026

View reviewed changes

examples/llm_qat/README.md Show resolved Hide resolved

realAsma reviewed Jan 7, 2026

View reviewed changes

modelopt/torch/quantization/plugins/transformers_trainer.py Outdated Show resolved Hide resolved

Address review

ace662b

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

realAsma reviewed Jan 8, 2026

View reviewed changes

modelopt/torch/quantization/plugins/transformers_trainer.py Outdated Show resolved Hide resolved

Remove QADTrainer.save_model()

6f82998

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

realAsma reviewed Jan 9, 2026

View reviewed changes

modelopt/torch/quantization/plugins/transformers_trainer.py Outdated Show resolved Hide resolved

realAsma reviewed Jan 9, 2026

View reviewed changes

modelopt/torch/quantization/plugins/transformers_trainer.py Outdated Show resolved Hide resolved

realAsma reviewed Jan 9, 2026

View reviewed changes

modelopt/torch/quantization/plugins/transformers_trainer.py Outdated Show resolved Hide resolved

realAsma reviewed Jan 9, 2026

View reviewed changes

modelopt/torch/quantization/plugins/transformers_trainer.py Outdated Show resolved Hide resolved

Slim down redundancies and optimize memory usage

cd6124f

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

realAsma approved these changes Jan 9, 2026

View reviewed changes

AAnoosheh enabled auto-merge (squash) January 9, 2026 23:59

Swap KDSFTTrainer inheritance

7079dee

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

AAnoosheh merged commit 5104513 into main Jan 10, 2026
35 checks passed

AAnoosheh deleted the aanoosheh/kd-trainer-streamline branch January 10, 2026 01:28

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Streamline KD & QAD transformers Trainers #708

Streamline KD & QAD transformers Trainers #708

AAnoosheh commented Dec 18, 2025

Uh oh!

copy-pr-bot bot commented Dec 18, 2025

Uh oh!

codecov bot commented Dec 18, 2025 •

edited

Loading

Uh oh!

kevalmorabia97 Dec 19, 2025

Uh oh!

AAnoosheh Dec 19, 2025

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

realAsma Jan 9, 2026

Uh oh!

Uh oh!

realAsma left a comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

Streamline KD & QAD transformers Trainers #708

Streamline KD & QAD transformers Trainers #708

Conversation

AAnoosheh commented Dec 18, 2025

What does this PR do?

Usage

Testing

Before your PR is "Ready for review"

Additional Information

Uh oh!

copy-pr-bot bot commented Dec 18, 2025

Uh oh!

codecov bot commented Dec 18, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Codecov Report

Uh oh!

kevalmorabia97 Dec 19, 2025

Choose a reason for hiding this comment

Uh oh!

AAnoosheh Dec 19, 2025

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

realAsma Jan 9, 2026

Choose a reason for hiding this comment

Uh oh!

Uh oh!

realAsma left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

codecov bot commented Dec 18, 2025 •

edited

Loading