RTL & FPGA Design · AI Accelerators

Syed Mohsin Shah

 

Graduate researcher building AI accelerators for ASIC and FPGA platforms. I design system-level RTL architectures and reusable digital hardware blocks, and build automated verification environments to validate design behavior.

Syed Mohsin Shah
About

What I work on

I focus on designing microarchitectures, integrating subsystems, ensuring timing closure, and verifying AI accelerators. My work spans accross performance-optimized designs, precision-aware arithmetic, and verification methodologies.

Toolbox

Skills

Languages

SystemVerilogVerilogVHDLC/C++PythonTclBashMATLAB

EDA Tools

Synopsys VCSDesign CompilerQuestaSimCadence VirtuosoXilinx VivadoIntel Quartus

DV Concepts

UVMSelf-Checking TestbenchesCDCSTATiming ClosureFunctional Coverage

Architecture

AXI / SPI / I2CDMANoCFSM ControlFault ToleranceDFT / BIST

Workflow

GitLinuxDockerMakefile
Journey

Experience

Research Assistant
University of Saskatchewan — Ko Lab
Sep 2024 – Present
  • Architected a 55-module FPGA SoC accelerator (Winograd PE array, hierarchical NoC, DMA engine, multi-bank memory, control FSMs) with end-to-end RTL design and subsystem integration.
  • Designed distributed SystemVerilog RTL and FSM control for compute scheduling and memory sequencing, integrating vendor DDR4 EMIF and BRAM IP over AXI.
  • Achieved 250 MHz timing closure on Intel Arria-10 via synthesis, STA, pipelining, retiming, and register balancing.
  • Reduced ASIC power 4–7% with variable-precision arithmetic units minimizing switching activity under Design Compiler.
  • Built self-checking testbenches with golden-reference models and Python/Tcl regression across 15+ configs, cutting debug turnaround 60%.
  • Evaluated spatial, depth-wise, and capacity-based tiling strategies for PE-array mapping, balancing compute utilization against on-chip memory and DSP constraints.
Research Collaboration — Edge AI Transportation (WIM)
Quarterhill
2026
  • Developing an embedded edge-AI FPGA platform for Weigh-In-Motion calibration, fusing computer vision and axle-sensor streams for real-time vehicle weight estimation.
  • Building automated multimodal data-acquisition and validation pipelines across 10,000+ test passages, profiling latency and accuracy pre-deployment.
Research Assistant
System-on-Chip Lab, NUST
Mar 2024 – Aug 2024
  • Researched fault-tolerant system architectures for deep neural networks, targeting resilience under hardware-induced errors.
  • Implemented TinyML inference pipelines for edge deployment, optimizing for latency under constrained resources.
  • Programmed FPGAs in Verilog for hardware customization, tuning performance and efficiency.
FPGA Intern
DreamBig Semiconductors
Jan 2024 – Mar 2024
  • Shadowed the FPGA/ASIC design team on-site to study industry RTL flows, verification methodology, and chip design practices outside coursework.
Research Assistant
Secured IoT Devices Lab, Peshawar
Oct 2022 – Jul 2023
  • Developed a LoRaWAN-based monitoring system with cloud integration, implementing embedded C firmware for sensor data acquisition on ESP32 nodes over SPI/I2C.
  • Built a Raspberry Pi–based gateway and secured the communication stack with cybersecurity protocols for a final-year IoT research project.
Selected Work

Projects

Add image

Winograd CNN Accelerator

RTL accelerator on ZCU102 hitting 5.3 TOPS on VGG16 via streamed dataflow optimization and simulation-based functional verification. Extended into an overlap-aware streaming architecture, ready for journal submission.

RTLFPGAZCU102
Add image

Power-Efficient POSIT Multiplier

Timing-aware Verilog RTL for decode, multiply, and recode pipeline stages; synthesized in Synopsys Design Compiler, validating 4–7% power reduction via switching-activity analysis.

ArithmeticLow Power
Add image

Depthwise Separable CNN Accelerator

Parallel RTL architecture for MobileNet-class models achieving ~20× speedup through dataflow parallelism.

RTLParallelism
Add image

DFT for Systolic Arrays

Scan chains and BIST for systolic PE arrays, improving fault coverage and enabling structured post-silicon validation.

DFTBISTVerification
Add image

Fault Tolerance Enhancement for DNNs

Investigated fault-tolerant architectures for deep neural network accelerators under hardware-induced error models.

ReliabilityDNN Hardware
Add image

Dynamic 32×64 LED Matrix Display

Xilinx FPGA design driving a high-resolution LED matrix at 30 fps, with UART-based visual data ingestion and clock/data synchronization.

RTLFPGAUART
Background

Education

M.Sc. Electrical Engineering

University of Saskatchewan
July 2026 · Deep Learning Architectures, Computer Architecture, EDA Simulation

B.Sc. Computer Systems Engineering

University of Engineering & Technology, Peshawar
2023
Beyond the Bench

Soft Skills

Leadership & Mentorship

Led a project team and mentored peers (7-member team, K2X); mentored students as a Teaching Assistant.

Communication & Presentation

Presented extensively in graduate courses; liaised with international clients and stakeholders.

Teaching & Community

TA for Probability & Statistics (UofS, 2026) and high-school math; Microsoft Learn Student Ambassador; IEEE Student Member.

Volunteering

Student Wellness Center, UofS — event logistics; Media Team, University Computer Society.

Customer Service

POS operations, cash handling, and client-facing service, APS Stationary.

Cross-Functional Collaboration

3+ years of general software engineering experience (K2X, Dcube, Auxcube, The Solutioners) coordinating with product, client, and research teams.

Contact

Let's build something.

Open to RTL / FPGA / verification roles and research collaborations. Reach out — I reply fast.

© 2026 Syed Mohsin Shah · Saskatoon, SK