Kai Yao · Selected writing

Blogs

Technical writing on machine-learning systems, developer tooling, and trustworthy AI. This archive begins with my work on Neural Coder at Intel and will grow with future writing across research and engineering.

At Intel, I helped develop Neural Coder, a component of Intel Neural Compressor that automated the insertion and evaluation of deep-learning optimizations in existing Python code. It combined static program analysis—syntax analysis, type inference, and call-graph parsing—with automated benchmarking and brought the workflow from PyTorch scripts to Hugging Face Transformers, Alibaba Cloud PAI-DSW, JupyterLab, and Visual Studio Code.

Neural Coder was demonstrated in Intel CEO Pat Gelsinger’s Intel Innovation 2022 keynote. This work was recognized internally at Intel with the AIA Division Achievement Award for Neural Coder innovation (2022) and the CESG SW AI Division Recognition Award for the Alibaba Cloud collaboration (2023); both are listed on my professional profile.

10×+
Intel-reported ResNet50 inference gain in the 4th Gen Xeon keynote demo, with accuracy maintained
18
Trainer-based PyTorch tasks supported by the Hugging Face integration
~1.7×
reported acceleration on 3rd Gen Xeon, typically with less than 1% accuracy loss
2
internal Intel division awards for Neural Coder innovation and the Alibaba Cloud collaboration

2022

4 articles · newest first
  1. Model optimization

    One-Click Acceleration of Hugging Face Transformers with Intel’s Neural Coder

    Showed how Neural Coder automatically enabled Optimum Intel INT8 quantization for Hugging Face Transformers. The evaluated workflow covered 18 Trainer-based PyTorch tasks and reported about 1.7× inference acceleration on 3rd Gen Intel Xeon processors, typically with less than 1% accuracy loss.

    Authors: Kai Yao, Yuwei Huang, Yu Liao, Wenjiao Yue, Haihao Shen, and Huma Abidi

  2. Developer tooling

    One-Click Quantization of Deep Learning Models with the Neural Coder Extension

    Presented a Visual Studio Code extension that applied static or dynamic INT8 quantization and BF16 optimization, preserved generated code changes as inspectable patches, and benchmarked candidate transformations directly inside the editor.

    Authors: Wenjiao Yue, Kai Yao, Suyue Chen, Haihao Shen, and Huma Abidi

  3. Industry collaboration

    Alibaba Cloud and Intel Neural Compressor Deliver Better Productivity for PyTorch Users

    Documented an Intel–Alibaba integration that brought Neural Coder into Alibaba Cloud PAI-DSW and exposed BladeDISC as an optimization backend, enabling one-click optimization and automatic before-and-after benchmarking from JupyterLab.

    Authors: Shen Li (Alibaba); Kai Yao, Shan Zhou, Haihao Shen, and Huma Abidi (Intel)

  4. ML systems

    One-Click Enabling of Intel Neural Compressor Features in PyTorch Scripts

    Introduced Neural Coder, a one-click workflow combining deep-learning optimization rules with Python syntax analysis, static type inference, and call-graph parsing to insert quantization code into PyTorch scripts and benchmark applicable optimizations.

    Authors: Kai Yao, Haihao Shen, and Huma Abidi