Programmable Compute-in-Transit using Integrated Photonics

by Imon Kundu
Director of Photonics at Optalysys
Modern hardware designs for AI and cryptography treat data transit and processing separately. At Optalysys we have demonstrated programmable Photonic hardware that computes functions on data that is in transit, enabling tens of GFLOPs of operations on data links.
The data movement bottleneck in AI
Data movement is increasingly the dominant bottleneck in modern computing systems, particularly for demanding workloads such as AI, Post-Quantum Cryptography (PQC), and Fully Homomorphic Encryption (FHE).
Conventional architectures treat computation and communication as separate processes, with
data transported to discrete processing units before operations can be applied. As system scale increases, this separation leads to inefficiencies in latency, bandwidth, and energy.
Conventional architectures treat computation and communication as separate processes, with data transported to discrete processing units before operations can be applied. As system scale increases, this separation leads to inefficiencies in latency, bandwidth, and energy.
Compute-in-Transit offers an alternative model, where computation is performed during data movement by embedding operations directly along the data path, shifting the focus from moving data to compute towards applying computation as data propagates through the system.
Scaling data movement-bound workloads
These workloads are characterised by extremely high computational demands, often reaching tera-scale to exa-scale operations per second. In AI systems, multiply-accumulate (MAC) operations on floating-point data dominate accelerator design, with modern hardware adopting reduced-precision formats to improve efficiency and reduce memory footprint.
In contrast, cryptographic workloads such as FHE rely on structured polynomial transformations, including Fast Fourier Transforms (FFT) and Number Theoretic Transforms (NTTs), which form the computational backbone of secure processing. Although these domains differ in structure, both involve repeated application of regular operations over large data sets, making performance increasingly sensitive to data movement and memory access patterns. This is further amplified by the creation and propagation of intermediate data during computation, which must be repeatedly stored, transferred, and transformed across processing stages.
Digital Signal Processing (DSP) cores and Application-Specific Integrated Circuits (ASICs) are widely used to accelerate the dominant operations in these workloads, including MAC operations in AI and large-scale FFT and NTT computations in FHE.
However, scaling to very large transform sizes—such as 65,536-point NTTs required in certain FHE workloads or even larger ones in some Zero-Knowledge Proof schemes —remains constrained by memory access patterns, interconnect complexity, and the cost of moving data between processing elements. In practice, these constraints limit overall system performance and highlight the growing imbalance between compute capability and data movement efficiency.
Compute-in-Transit – adding compute capability directly along the data path
Conventional architectures treat data movement and computation as distinct stages, requiring data to be transferred across memory hierarchies and interconnects before processing. As workloads scale, this repeated movement of data, including intermediate results, introduces increasing overheads and limits achievable performance. Recent work has started to explore tighter integration between computation and data movement, including hardware–software co-design strategies and photonic implementations for structured workloads, providing a foundation for architectures that more closely align computation with dataflow.
In this paper, we present a prototype demonstration of the Compute-in-Transit architecture using a Photonic Integrated Circuit (PIC), enabling computation to be applied directly along the data path during transmission. The proposed implementation supports a range of mathematical operations through a programmable serial compute pipeline operating across high-speed transceiver lanes in a standard QSFP56 port. We demonstrate a prototype hardware platform with a custom driver IP that enables synchronous operation over asynchronous links. The prototype achieves sustained operation through integrated self-link fault detection and fully automated photonic calibration mechanisms.
At Optalysys we’re pioneering the architectural revolution enabling photonic compute-in-transit. Get in touch with us to find out how we can bring efficiency gains to your use case →
Want more like this delivered straight to your inbox?
Subscribe to stay in the loop

