# Task You are an expert CUDA, TensorRT, and C++ ...
बनाया गया: 17 जुलाई 2026
का उपयोग करके उत्तर दिया गया GPT-5.6 Thinking द्वारा Chat01
बनाया गया: 17 जुलाई 2026
का उपयोग करके उत्तर दिया गया GPT-5.6 Thinking द्वारा Chat01
You are an expert CUDA, TensorRT, and C++ performance engineer.
Your goal is to maximize the speed of my YOLO TensorRT library while keeping the same detection accuracy.
Current performance:
This is unacceptable.
Move every possible operation from CPU to GPU.
Assume I have a powerful NVIDIA GPU with enough VRAM.
Accuracy must not decrease.
Especially optimize:
video_processor.cuAvoid unnecessary CPU work, memory copies, allocations, and CPU-GPU synchronization.
Reuse GPU buffers whenever possible.
Use asynchronous CUDA APIs whenever possible.
The library uses two TensorRT models:
models_modelIf there are no dependencies between them, execute them in parallel using separate CUDA streams and separate TensorRT execution contexts to maximize GPU utilization.
Also pipeline frame processing so while one frame is running inference, other frames can already be preprocessing or postprocessing.
Frames are completely independent, so you may process multiple frames simultaneously if it improves throughput.
Fix any bugs, crashes, memory leaks, race conditions, or performance bottlenecks you find during optimization.
Do not explain anything.
Do not add comments.
Do not summarize your changes.
Only return the complete edited source files that were modified.
pythonimport cv2, numpy as np, math, random, os, sys, statistics print(cv2.__version__)
pythonimport numpy as np, cv2, math def custom(proto, coeff, crop, out_size, box): # proto [32,h,w], crop x1,y1,x2,y2, out W,H, box x,y,w,h _, mh, mw = proto.shape sig = 1/(1+np.exp(-np.tensordot(coeff, proto, axes=(0,0)))) cx1,cy1,cx2,cy2=crop cw,ch=cx2-cx1,cy2-cy1 ow,oh=out_size bx,by,bw,bh=box out=np.zeros((bh,bw),np.uint8) for y in range(bh): oy=by+y sy=(oy+0.5)*ch/oh-0.5+cy1 yb=math.floor(sy); ay=sy-yb y0=max(cy1,min(yb,cy2-1)); y1=max(cy1,min(yb+1,cy2-1)) for x in range(bw): ox=bx+x sx=(ox+0.5)*cw/ow-0.5+cx1 xb=math.floor(sx); ax=sx-xb x0=max(cx1,min(xb,cx2-1)); x1=max(cx1,min(xb+1,cx2-1)) v00=sig[y0,x0];v01=sig[y0,x1];v10=sig[y1,x0];v11=sig[y1,x1] v=(v00+(v01-v00)*ax)+((v10+(v11-v10)*ax)-(v00+(v01-v00)*ax))*ay out[y,x]=255 if v>0.5 else 0 return out rng=np.random.default_rng(1) for k in range(20): mh,mw=160,160 proto=rng.normal(size=(32,mh,mw)).astype(np.float32) coeff=rng.normal(size=32).astype(np.float32) cx1,cy1,cx2,cy2=0,10,160,150 ow,oh=640,480 bx,by,bw,bh=37,21,202,113 logit=np.tensordot(coeff,proto,axes=(0,0)).astype(np.float32) sig=(1/(1+np.exp(-logit))).astype(np.float32) resized=cv2.resize(sig[cy1:cy2,cx1:cx2],(ow,oh),interpolation=cv2.INTER_LINEAR) ref=(resized[by:by+bh,bx:bx+bw]>0.5).astype(np.uint8)*255 got=custom(proto,coeff,(cx1,cy1,cx2,cy2),(ow,oh),(bx,by,bw,bh)) diff=np.count_nonzero(ref!=got) print(diff) if diff: break
[include/yolos/core/cuda_preprocessing.cu]
include/yolos/core/cuda_preprocessing.hpp
[include/yolos/core/trt_session_base.hpp]
include/yolos/core/trt_utils.hpp
[include/yolos/tasks/segmentation.hpp]
include/yolos/tasks/sahi_fixed_seg.hpp
[include/yolos/tasks/video_processor.cu]