Services & research

Neuromorphic and low-power edge AI, benchmarked rather than assumed

Inference on a power budget: spiking networks on neuromorphic accelerators, and quantised models on conventional edge hardware. We measure with instruments and report accuracy beside energy.

Built and validated under NAVIR, funded through the dAIEDGE network.

Bring a model and a budget

An engineer, not a sales rep, replies within one business day.

We reply within one business day. Your details are used only to answer your enquiry. Privacy.

Three problems that show up on a power budget

Inference on a device that cannot be plugged in, cooled or easily reached is a different engineering problem from inference in a data centre. These are the three constraints that decide most projects.

The model does not fit the power budget

A model that runs happily on a workstation has an energy cost per inference that, multiplied by a duty cycle and a battery, makes the product impossible. This is usually discovered late, after the model is built and the enclosure specified, when every remaining option is a bad one.

Efficiency claims cannot be compared

Most published energy figures are derived from operation counts rather than measured, and the measured ones rarely state their scope. Whole system or one component, idle subtracted or not, which model, how long a run. Two honest sources can differ by an order of magnitude, which leaves you unable to choose between platforms on the evidence available.

The device cannot be cooled, and you cannot get to it

Sealed enclosures, no fan, and a site visit that costs more than the hardware. Thermal throttling turns a benchmark figure into a fiction, and any approach needing regular physical access is not viable however well it performs on a bench.

What we build

Models that run inside a power budget, on hardware that cannot be cooled, charged or connected the way a data centre can. Three kinds of work.

  • Spiking models for neuromorphic accelerators. Architectures designed around a constrained operator set rather than ported to it, with conversion and quantisation-aware training for on-chip deployment.
  • Optimising conventional models for edge hardware. Quantisation, structured pruning and distillation, measured on the target architecture. Mature, production-proven, and where most projects should begin.
  • Multimodal fusion at the edge. Combining vision with audio or other sensor channels so a system degrades gracefully when one input becomes unreliable, rather than collapsing.

Under NAVIR we built a real-time audio-visual speech recognition system implemented as spiking neural networks on a neuromorphic accelerator, letting an operator give voice commands to a robot arm in an environment too noisy for conventional speech recognition. The whole system runs on a single-board computer with the accelerator attached. By the project's own review of the published field, it is one of the first demonstrations of audio-visual speech recognition on this class of processor.

This is research-level work and we describe it that way deliberately. It is a validated demonstrator, not a product, and it has not been operated across a fleet or over months. What is transferable today is the engineering method and the benchmarking discipline, and we are actively looking for the next project to take it further.

Why the results make us want to do more of this

On our video model the accelerator sustained 14.55 inferences per second, comfortably above real time for the task, while drawing roughly an order of magnitude less energy per inference than conventional backends on the same board. The audio-video model ran at 2.61 inferences per second, still ample for speech, at well under a tenth of the CPU's per-inference energy.

Getting throughput and efficiency together, rather than trading one for the other, is what makes this interesting. For a device on a battery or harvested power, in a sealed enclosure with no fan, that combination decides whether a product is possible at all. The applications we find most compelling are the ones nobody currently attempts because the power budget rules them out.

Those figures are measurements of our models in our configuration, on one workload. They are not a benchmark of any product and we would not present them as one. Even our own two models behaved differently on the same board, which is precisely why the next section exists.

Benchmarking is the service

Most energy claims in this field are derived from operation counts rather than observed. We measure the supply with a meter, subtract and publish idle draw, report throughput alongside energy, and state the scope of every figure: whole system or one component, same instrument or two, how long the run was.

That protocol is the reusable output of our research, and it answers the only question that actually matters when choosing a platform, which is what your model costs on your target and what task performance comes with it. The full method is written up here.

Where this is worth investigating

Strong candidates

  • Battery, solar or harvested power
  • Continuous rather than occasional inference
  • Sealed or passively cooled enclosures
  • Always-on sensing, keyword spotting, event detection
  • Products that a conventional power budget rules out entirely
  • Research and prototype work inside a funded consortium

Benchmark conventional optimisation first when

  • Mains power is available and energy is not binding
  • An existing model must ship without redesign
  • The task leans heavily on long-range temporal context
  • High frame-rate vision is the workload
  • A production deployment is needed on a fixed date

The right column is not an argument against the technology. It is an argument for establishing the baseline first, so that whatever you choose, you chose it on evidence.

Questions we get asked

How mature is your neuromorphic work?

Research level. We built and validated a working demonstrator inside an EU-funded project, on a bench and on a robot arm, and we have not operated it across a fleet or over months. We would take on a research or prototype engagement today and we would tell you plainly that a production deployment is a further step.

What energy saving should we expect?

On our own models we measured roughly an order of magnitude less energy per inference than conventional backends on the same board, at higher throughput. Those are our models on one workload, and even our own two models behaved differently on the same hardware. The only number that will predict yours is a measurement of yours.

What accuracy do spiking models reach?

It depends entirely on how much your task needs operators that a given accelerator executes natively. Our constrained-vocabulary command model reached 98.6% command accuracy under noise. Our open-vocabulary lip-reading component reached 34% word error rate, where unconstrained research models reach around 10%. Both numbers are real and they measure different problems.

Do we have to use neuromorphic hardware to get low power?

Not necessarily, and we would benchmark both routes. Quantisation, structured pruning and distillation on conventional edge hardware are mature, production-proven and often sufficient. Neuromorphic accelerators open a different envelope, and which one fits is an empirical question rather than a matter of opinion.

Can you join a consortium on this?

Yes, and this is where most of our neuromorphic work happens. We delivered an audio-visual spiking pipeline inside an EU-funded project and can take a technical work package on embedded ML, edge inference or benchmarking methodology.

Next step

Bring a model and a power budget

The two-week audit measures your model on your target hardware, conventional and neuromorphic where both apply, and reports energy, throughput and task performance together. Written findings at the end.

Free 45-minute call. Then a two-week data audit with a written go/no-go before you commit to anything.

Send us the problem

An engineer, not a sales rep, replies within one business day.

We reply within one business day. Your details are used only to answer your enquiry. Privacy.