ESSCIRC 2023 · JxCDC 2024 · JSSC 2025

Making MRAM-based in-memory computing more reliable.

This project follows one 22 nm MRAM in-memory-computing macro from a parasitic-aware behavioral model to measured silicon. Statistical error compensation, mixed-signal design, tapeout, and PYNQ bring-up are the stages in between.

Saion K. Roy · Han-Mo Ou · Mostafa G. Ahmed · Peter Deaville · Bonan Zhang · Naveen Verma · Pavan K. Hanumolu · Naresh R. Shanbhag

SEC-enabled MRAM IMCStructured array correction
Complete figure showing the annotated die micrograph, macro area distribution, and measured performance summary of the 22 nanometer MRAM in-memory-computing prototype
512 × 512 MRAM array · 128 ADC columnsCommercial 22 nm FD-SOI
22 nmCommercial FD-SOI measured prototype
512 × 512Physical 1T1R MRAM array
128Parallel 6-bit ADC columns
2.7–6 dBMeasured SEC SNDR improvement
5×Energy reduction at iso-SNDR

The complete research arc

One idea, carried from behavioral model to measured silicon.

The project begins with one physical observation: row position changes analog contribution. Everything downstream follows from it. Each stage turns the output of the previous stage into a concrete design decision.

01 · Modeling

Describe the error with α\alpha

MTJ variation, BL/SL parasitics, read noise, and ADC conversion pull the analog output away from ideal arithmetic. αij\alpha_{ij} captures the contribution of row ii as observed by column jj, making the spatial structure explicit.

02 · Design

Correct it with γ\gamma and θ\theta

γi\gamma_i pre-scales each row and θj\theta_j normalizes each output column so that θjγiαij≈1\theta_j\gamma_i\alpha_{ij}\approx 1 statistically. Precision studies fix the fixed-point path, and OCCS, the array, ADCs, and control close into one macro.

03 · Tapeout

Close the design against one identifier

Frozen interfaces and golden vectors carry through mixed-signal verification, physical and package closure, and a release manifest the bench can execute.

04 · Testing

Measure compute SNDR on silicon

Python, a PYNQ-Z2, and a custom PCB drive the packaged macro. Per-column calibration and code-conditioned sampling turn raw ADC captures into the measured result.

Research thesisSEC works because the dominant array error has spatial structure that can be learned with one correction per row and one normalization per column.

Why the problem is hard

A useful signal can be a fraction of one percent.

For a 128-row dot product, the paper’s representative conductance step is ΔGS=80 μS\Delta G_{\mathrm S}=80\,\mu\mathrm{S} against a nominal column conductance GS=20 mSG_{\mathrm S}=20\,\mathrm{mS}. Bitline and sourceline voltage gradients, current-sensor noise, column mismatch, and supply drop compete directly with that small step.

ΔGSGS=80 μS20 mS=4×10−3=0.4%\frac{\Delta G_{\mathrm S}}{G_{\mathrm S}}=\frac{80\,\mu\mathrm{S}}{20\,\mathrm{mS}}=4\times10^{-3}=0.4\%
Parallelism compresses the analog margin.

A larger dot product activates more rows and accumulates more current, while the step between adjacent ideal outputs remains small. Compute SNDR becomes the link between array scaling and useful arithmetic.

Inputx\mathbf{x} and w\mathbf{w} define the ideal dot product

The experiment selects an activation and weight state with a known ideal code.

Row scale γi\gamma_iPre-scale each activated row

A learned 7-bit factor adjusts the digital input before array evaluation.

Array response αij\alpha_{ij}Position shapes analog contribution

MRAM state, BL/SL parasitics, and row location determine the column current.

OCCS + ADCSense and digitize the current

Offset-compensated sensing feeds a 6-bit SAR conversion for each ADC column.

Calibrate + scoreNormalize output and compute SNDR

Per-column affine calibration and code-conditioned error turn raw codes into the reported metric.

Signal pathThe same chain explains both the architecture and the experiment: γi\gamma_i acts before the physical array, OCCS preserves the small analog step, and column calibration makes measured codes comparable to the ideal output.

Prototype at a glance

Four levels describe the same dot product.

A physical 1T1R array stores MRAM states, differential 2T2R pairs encode signed weights, OCCS and the SAR ADC convert column current, and SEC scales rows and columns to improve bank-level compute SNDR.

DEVICE

Physical memory

Commercial 22 nm FD-SOI, 512×512512\times512 physical 1T1R MRAM cells, operated at 0.8 V0.8\,\mathrm{V}.

LOGIC

Signed weight

A differential 2T2R logical representation forms signed 4-bit weights across binary-weighted physical columns.

CIRCUIT

Column conversion

128 ADC columns combine offset-compensated current sensing, a 6-bit SAR ADC, and local control.

ALGORITHM

SEC correction

Learned row scaling and per-column normalization improve task-agnostic bank-level compute SNDR.

Paper figure showing the OCCS and SEC-enabled MRAM in-memory-computing macro architecture
Integrated architecture. Physical array, binary-weighted columns, OCCS, ADC conversion, SEC scaling, and local control are treated as one signal path.

Engineering record

From mathematical intent to a repeatable measurement.

This repository documents the decision gates that papers usually compress: what was modeled, what was frozen for implementation, what was checked at tapeout, what the board had to make observable, and how the measurement state space was sampled.

DESIGN

Behavior → architecture

Trace α\alpha, γ\gamma, and θ\theta from the behavioral model into OCCS, fixed-point precision, and dataflow.

Open design principles →
IMPLEMENT

Netlist → packaged die

Follow the mixed-signal design gates from bit-accurate correlation through physical verification, handoff, packaging, and post-silicon readiness.

Open tapeout workflow →
TEST

Host → calibrated result

Follow the PYNQ, PCB, scan-chain, calibration, sampling, and analysis sequence used to interrogate silicon.

Open test platform →

The role of SEC

The array learns its own correction.

Statistical error compensation calibrates the physical compute bank. The application weights stay fixed while SEC learns row factors γi\gamma_i and column normalizers θj\theta_j from the measured hardware response.

This raises task-agnostic bank-level compute SNDR. The CIFAR-10 experiment then uses the measured final fully connected layer of ResNet-20 to show how that signal-quality change affects classification.

From SNR to SNDR

The ESSCIRC 2023 paper introduces the result as an SNR boost. The JSSC 2025 extension separates temporal noise and state-dependent distortion through the broader compute-SNDR methodology used throughout this site.