pytorch 相关面试题

Creado el: 2 de septiembre de 2025

Respondido usando GPT-5 Thinking por Chat01

Pregunta

pytorch 相关面试题

Pregunta

pytorch 相关面试题

Pensamiento

Creating PyTorch interview questions

The user seems to want a comprehensive list of PyTorch-related interview questions. I'll provide a set categorized into different sections: basics, tensors, autograd, nn, optimization, etc. Questions will cover things like model behavior, memory management, and specific topics like mixed precision, dynamic vs static graphs, and CUDA. This should be a thorough list!

Designing interview questions for PyTorch

The user wants me to create a set of PyTorch interview questions, categorized by topics like distributed training, mixed precision, quantization, pruning, and JIT vs TorchScript. I'll include practical tasks, like writing a custom DataLoader, and provide concise answers, code examples, and key hints. Keeping things clear and practical!

Including troubleshooting questions

I want to include advanced troubleshooting questions about CUDA OOM debugging and deterministic behavior. I'll mention settings like torch.backends.cudnn.deterministic=True, and in PyTorch 2.0, torch.use_deterministic_algorithms(True) for reproducibility. Also, I’ll touch on matmul precision and relevant environment variables like CUBLAS_WORKSPACE_CONFIG.

Clarifying TF32 and precision settings

I need to carefully explain TF32 and how it's controlled in PyTorch through options like torch.backends.cuda.matmul.allow_tf32. For torch.set_float32_matmul_precision, the "high" setting means higher performance with less precision, while "medium" allows for TF32. It's important to highlight these details and also mention inference optimizations like torch.inference_mode().

Thought for 57s

Respuesta

太好了!下面给你一份 PyTorch 相关面试题清单(含答案要点与常用代码片段)。按主题分区,既能快速刷题,也可当复习提纲。

一、张量 & 基础

  1. tensor.view() 和 reshape() 有什么区别?
    要点:view 仅在内存连续(is_contiguous())时零拷贝;不连续需 contiguous();reshape 可能返回拷贝。

  2. 广播机制如何工作?何时会报错?
    要点:从尾维对齐;维度相等或其一为 1 可扩展;不满足则形状不兼容。

  3. in-place 操作(带 _)对 Autograd 的影响?
    要点:可能覆盖计算图中需要的中间值 -> 抛错;对叶子张量(requires_grad=True)也要谨慎。

  4. detach()、item()、with torch.no_grad() 区别?
    要点:detach 返回共享内存但无梯度的张量;item() 将标量张量转为 Python 数字(会断梯度);no_grad 关闭梯度记录(评估/推理)。

  5. 解释 dtype、device、pin_memory 的常见搭配。
    要点:tensor.to(device, non_blocking=True);DataLoader pin_memory=True 加速 H2D;AMP 时常配合。

二、Autograd 机制

  1. 计算图是如何构建与释放的?
    要点:正向时记录算子与依赖;tensor.grad_fn;默认反向后释放图,需二阶梯度用 backward(retain_graph=True) 或对输出 create_graph=True。

  2. 多次反向/梯度累积如何写?
    要点:loss/accum_steps 再 backward();每 N step 执行 optimizer.step(); optimizer.zero_grad(set_to_none=True)。

  3. 自定义反向:torch.autograd.Function 如何用?
    要点:save_for_backward 存中间量;backward(ctx, grad_output) 返回与输入同数目的梯度。

三、nn.Module 与参数管理

  1. model.train() vs model.eval() 有何影响?
    要点:影响 Dropout/BN 等行为与统计;不影响 requires_grad。

  2. state_dict 的保存/加载与坑?
    要点:推荐存/载 state_dict 而非整模型;map_location;DataParallel 会多出 module. 前缀(需去掉)。

  3. 冻结/解冻部分层参数怎么做?
    要点:for p in module.parameters(): p.requires_grad=False;优化器只传可训练参数。

四、优化器 & LR 调度

  1. Adam 与 AdamW 的区别?
    要点:AdamW 采用解耦权重衰减;常对 bias/LayerNorm/BN 不施加衰减(参数组过滤)。

  2. ReduceLROnPlateau 与其他调度器的用法差异?
    要点:前者需传入度量 scheduler.step(val);其他通常每 step/epoch 调一次。

  3. 学习率预热(warmup)如何实现?
    要点:LambdaLR 自定义函数 or 组合 Cosine;示例:线性 warmup + cosine。

python
def lr_lambda(step): if step < warmup: return step / warmup t = (step - warmup) / (total - warmup) return 0.5*(1+math.cos(math.pi*t)) scheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda)
  1. 梯度裁剪(clip)何时使用?
    要点:RNN/大模型易梯度爆炸;torch.nn.utils.clip_grad_norm_。

五、DataLoader & Dataset

  1. 写一个自定义 Dataset 需要实现什么?
    要点:__len__、__getitem__;必要时自定义 collate_fn 处理变长数据。

  2. num_workers、persistent_workers、prefetch_factor 如何影响性能?
    要点:更多 worker 并行加载;长训练可设 persistent_workers=True;充分利用 CPU/I/O。

  3. 如何保证多进程数据加载的可复现?
    要点:设置全局种子;worker_init_fn 中基于 worker_id 再设子种子;使用 DistributedSampler 时设 set_epoch()。

六、GPU/性能加速

  1. AMP(自动混合精度)如何写?autocast + GradScaler 基本范式?
    要点:训练时:
python
scaler = torch.cuda.amp.GradScaler() for x,y in loader: optimizer.zero_grad(set_to_none=True) with torch.cuda.amp.autocast(): loss = criterion(model(x), y) scaler.scale(loss).backward() scaler.step(optimizer) scaler.update()

推理:只用 autocast() 或 torch.inference_mode()。

  1. bfloat16 与 float16 的差异?
    要点:bfloat16 动态范围大、训练更稳,通常不需要 GradScaler;硬件支持取决于 GPU/CPU。

  2. 提升吞吐的几个实践?
    要点:channels_last(model.to(memory_format=torch.channels_last))、pin_memory、non_blocking=True、合理的 batch、cudnn.benchmark=True(动态输入慎用)、避免 Python 端瓶颈。

  3. 避免 CUDA OOM 的技巧?
    要点:AMP、梯度累积、torch.utils.checkpoint、减少激活保存(删多余中间变量)、验证阶段 no_grad/inference_mode。

七、分布式训练(DDP)

  1. DataParallel 与 DistributedDataParallel 区别?
    要点:DP 单进程多线程,复制模型到多卡,性能差;DDP 多进程多卡,通信更高效(NCCL/Gloo),推荐。

  2. DDP 基本启动流程?
    要点:torch.distributed.init_process_group(backend='nccl'),每进程绑定 local_rank;用 DistributedSampler;在封装 DDP 前把模型/优化器初始化好。

  3. find_unused_parameters 什么时候开?
    要点:有条件分支导致部分参数不参与前向时;但会有性能损耗。

  4. DDP 同步/异步 BN 如何处理?
    要点:多机多卡用 SyncBatchNorm(convert_sync_batchnorm),注意开销。

八、部署 & TorchScript/ONNX/torch.compile

  1. torch.jit.trace vs torch.jit.script?
    要点:trace 只记录实际执行路径;对动态控制流不可靠;script 解析 TorchScript,支持控制流但语法有约束。

  2. 导出 ONNX 的基本示例与注意点?
    要点:torch.onnx.export(model, x, 'm.onnx', opset_version=17, dynamic_axes=...);部分算子/动态形状不友好;用 onnxruntime 做推理。

  3. PyTorch 2.x 的 torch.compile 怎么用?
    要点:model = torch.compile(model, mode='max-autotune');通常 先 compile 再 DDP;编译冷启动开销、动态图/稀疏/自定义算子可能降级或不支持。

九、量化/剪枝/蒸馏

  1. 动态量化 vs 静态(PTQ)vs QAT 区别?
    要点:动态量化仅权重 8bit、激活运行时量化(线性层常见);静态后训练量化需要校准;QAT 训练期插入伪量化,精度最好。

  2. PyTorch 量化入口在哪里?
    要点:torch.ao.quantization(或 quantization 子包);FX graph 模式量化更常用。

  3. 剪枝如何做?
    要点:torch.nn.utils.prune;非结构化 vs 结构化;剪后需 prune.remove 固化掩码。

  4. 知识蒸馏的关键实现?
    要点:教师软标签(温度 T)、学生交叉熵 + KL 损失的加权。

十、调试、日志与复现

  1. 如何完整复现一次训练?
    要点:固定随机种子(PyTorch/NumPy/随机库);设 torch.use_deterministic_algorithms(True);配置 CUDNN(deterministic=True、视需求关闭 benchmark);记录软件版本、超参、数据哈希。

  2. 常见的训练数值不稳定现象与定位?
    要点:loss NaN/Inf;打印梯度范数;检查学习率/梯度缩放/激活分布;逐步缩小最小复现。

  3. 日志与可视化?
    要点:TensorBoard、W&B;记录 LR、grad norm、显存、吞吐。

十一、实战小题(可现场写)

A. 自定义 Dataset+collate_fn(变长序列 Padding)

python
class MyDS(torch.utils.data.Dataset): def __init__(self, items): self.items = items def __len__(self): return len(self.items) def __getitem__(self, i): return torch.tensor(self.items[i]) def pad_collate(batch): lens = [x.size(0) for x in batch] maxlen = max(lens) out = torch.zeros(len(batch), maxlen, dtype=batch[0].dtype) for i,x in enumerate(batch): out[i,:x.size(0)] = x return out, torch.tensor(lens) loader = DataLoader(MyDS([[1,2],[3,4,5]]), batch_size=2, collate_fn=pad_collate)

B. 梯度检查点(节省显存)

python
from torch.utils.checkpoint import checkpoint class Block(nn.Module): def __init__(self): super().__init__(); self.net = nn.Sequential(...) def forward(self, x): return checkpoint(self.net, x)

C. 只对部分参数使用权重衰减(AdamW)

python
decay, no_decay = [], [] for n,p in model.named_parameters(): (no_decay if n.endswith('bias') or 'norm' in n.lower() else decay).append(p) optimizer = torch.optim.AdamW( [{'params': decay, 'weight_decay': 0.01}, {'params': no_decay, 'weight_decay': 0.0}], lr=3e-4 )

D. 分布式训练骨架(单机多卡)

python
def main(rank, world_size): torch.cuda.set_device(rank) dist.init_process_group('nccl', rank=rank, world_size=world_size) model = nn.parallel.DistributedDataParallel( build_model().to(rank), device_ids=[rank] ) train_loader = DataLoader(ds, sampler=DistributedSampler(ds)) ...

E. AMP 推理(更快更省显存)

python
model.eval() with torch.inference_mode(), torch.cuda.amp.autocast(): y = model(x)

加分题(常被问到的细节)

  • 为什么 optimizer.zero_grad(set_to_none=True) 更省显存/更快?
    要点:避免写入 0,后续第一次反向按需分配梯度。
  • BatchNorm 与小 batch 的稳定性问题?如何缓解?
    要点:用 GroupNorm/LayerNorm、SyncBN、多卡聚合、更大动量。
  • non_blocking=True 必须配合什么才有效?
    要点:DataLoader pin_memory=True。
  • 模型热启动/部分加载参数怎么做?
    要点:strict=False + 处理缺失/额外键。
  • 何时用 torch.inference_mode() 而不是 no_grad()?
    要点:全推理场景更快更省,因为不更新版本计数;训练/需要修改张量时别用。

需要我把这些题目整理成一份可打印的 PDF 或者生成一份带参考答案的题库(Markdown/Docx)吗?我也可以按“入门/中级/高阶”难度拆分给你。

Compartir este Q&A