{"slug": "functional-ultrasound-imaging-from-scratch", "title": "Functional Ultrasound Imaging from Scratch", "summary": "Functional ultrasound imaging (fUSI) can deliver up to 10 times higher linear resolution and 1000 times more voxels than functional MRI while requiring no magnet, according to a deep dive by Patrick Mineault, who published a series of Python tutorials on GitHub to make fUSI analysis accessible outside Matlab. The modality captures rapid movies of brain tissue by tracking hemodynamic blood-flow changes that follow neural activity, and it pairs with low-intensity focused ultrasound (LIFU) to modulate targeted brain regions as a treatment for neuropsychiatric and neurological disorders. fUSI has been demonstrated intraoperatively and in a patient with an acoustically transparent cranial implant, though ultrasound cannot easily penetrate an intact skull.", "body_md": "Functional ultrasound imaging (fUSI) is a brain imaging modality that is just starting to be demonstrated in human studies. I’m sure many will be familiar with s*tructural* ultrasound imaging, which forms static images using ultrasound, as in a sonogram[1](#footnote-1). *Functional* ultrasound instead captures *movies* of brain tissue very rapidly, tracking subtle changes caused by the blood flow that follows neural activity.\n\nWhile this is the same source of contrast as functional MRI, it has the potential to be ten times higher linear resolution (1000X the voxels), smaller (no magnet!), cheaper (no magnet!), and higher signal-to-noise. The modality naturally pairs with low-intensity *focused* ultrasound (LIFU, or confusingly, fUS) which focuses energy on a target area of the brain to modulate it, a new treatment option for neuropsychiatric and neurological disorders. Extensions leveraging contrast agents or gene therapies could go beyond hemodynamic responses to measure vasculature at super-resolution or even directly measure neural activity. And while ultrasound cannot easily penetrate the skull, fUSI has been demonstrated intraoperatively and in a patient that received an acoustically transparent cranial implant following reconstructive skull surgery.\n\nI did a deep dive into fUSI last year, and found that much of the technical information is hidden away in the methods sections of application papers, or spread throughout long books on ultrasound imaging as a whole. fUSI picks up on the hemodynamic response, but how? What’s the source of contrast? Does it measure blood flow or volume? How does it deal with motion of the head? What is ultrasound, anyway?\n\nI went down a rabbit hole—including reading [an 800-page textbook](https://www.sciencedirect.com/book/monograph/9780123964878/diagnostic-ultrasound-imaging-inside-out) pretty much cover-to-cover—so you don’t have to. I will break down how fUSI works physically and how it’s analyzed. This will give you a good basis to study fUSI computationally. Because much of the analysis toolkits available for fUSI sims are locked up in Matlab, [I’ve posted a series of Python tutorials on Github](https://github.com/patrickmineault/fus) so that you (or your agent) can follow along. At the end of this essay, you should understand what fUSI is, how it works, and why it is nowhere near its ceiling.\n\n*This is a much-expanded version of an article I posted on X in February. I wasn’t happy with the original, so I took a second pass: this piece includes new images and video, goes deeper into physics, and pinpoints where fUSI signal processing could be improved.*\n\n# How fUSI works at a high level\n\nFunctional ultrasound imaging, like conventional ultrasound imaging, uses transducers—which act as both transmitters and receivers—to communicate ultrasound through tissue. That ultrasound travels mostly unimpeded through brain tissue[2](#footnote-2), backscattering where it encounters a change in impedance, like a red blood cell, an air bubble, or the skull. By measuring these backscattered waves, it is possible to computationally reconstruct local changes in impedance. The resolution of the reconstructed images is proportional to the wavelength of the communicated ultrasound; that can mean up to ~100 um linear resolution when sonifying at 18 MHz, at the high end of what’s used clinically.\n\nThat explains how we can measure high-resolution structural images, but what about *functional* imaging? When neurons fire, they consume energy, which triggers vasodilation and a delayed influx of blood to the active region: hemodynamics. That blood has different properties than baseline: there’s more of it, it moves more rapidly, it has a different oxygen content, and it has a different color. Multiple brain imaging modalities take advantage of this fact, measuring changes in some or all of these properties, including fMRI (BOLD) and fNIRS. fUSI uses the fact that moving red blood cells cause shifting patterns of constructive and destructive interference in backscattered wavefronts, the fine-grained texture in ultrasound images known as speckle. fUSI acquires images of the brain very rapidly—over a kHz—and tracks changes in speckle over time. The amount of change in these images relates to blood flow and volume. The percentage signal change measured with fUSI is remarkably high: up to 20% compared to fMRI’s few-percent BOLD signal.\n\nHere’s an intuitive analogy[3](#footnote-3) for how that works. Imagine a nice high-res image of Brooklyn Bridge at night, taken with a long shutter speed. The surface of the water looks smooth and unbroken. But where is the movement in this image? To track movement, we can take a high-speed movie of the scene and measure how the brightness of each pixel varies in a window of time using the standard deviation. Color coding the resulting image, the shimmering surface of the water appears bright red, even though there may be no bulk flow, ignoring the tide. One could use this to quantify the water’s choppiness, or measure how water reacts to a boat speeding across it. Video analysis can thus reveal flow[4](#footnote-4).\n\nIf you know anything about photography and videography, I’m sure you can see the flaws in this analysis plan. If I crank up the frame rate and therefore shorten the exposure, won’t that decrease the SNR unacceptably? If I handhold the camera, won’t that mess up the analysis? Aren’t I discarding too much information by just calculating the standard deviation? Indeed, the same problems that plague video-based motion analysis also affect fUSI. Let’s see how these break down in detail.\n\n# A detailed breakdown of fUSI\n\n## Ultrafast plane waves and reconstruction\n\nConventional ultrasound forms images by scanning a focused beam line and sweeping it across the field of view, sonar-in-a-submarine style. fUSI uses a different, multiplexed strategy: instead of a focused beam, it transmits unfocused plane waves with flat wavefronts that insonify the whole imaging plane at once, so a single pulse can produce a whole image. Multiple detectors can then resolve the source of the backscatter. The repetition rate is limited by how much time it takes for the echo to come back to the transducer: for a 10 cm round trip and 1500 m/s propagation, the repetition rate reaches 15 kHz. In practice, plane waves transmitted at multiple angles and multiple repetitions are coherently compounded, decreasing the effective framerate to something manageable.\n\nThe reconstruction problem is then to infer the density of scatterers in the tissue given the time series for each element. To do that, we need to think about how ultrasound propagates through tissue, which can be done using physics**™**. It’s a fun exercise to derive the wave equation for sound in a fluid from scratch: it is quite similar, in fact, to deriving the Navier-Stokes equation, but with a different set of approximations. If you assume a homogeneous background medium, you can derive the propagation equations analytically, and you can simulate the equations using Fourier transforms.\n\nTo perform the reconstruction, you need to notice that if you measure a signal s(t) at time t on a given transducer, there is only a small slice of space that the signal could come from: that ambiguity takes the form of a conic section. Each element has a slightly different geometry with respect to the wavefront and the tissue, and the (conic) ambiguity resolves itself by compounding the conic sections. There are a few different ways to do this, with a popular one being the [delay-and-sum beamforming algorithm](https://arxiv.org/abs/2007.11960). For each pixel in the image, you calculate how long a sound wave takes to travel from the transducer to that point and back, apply those delays to the raw signals, and sum them up in image space. Constructive interference reveals the scatterer locations.\n\nAs is often the case for these things, you can instead cast this as a linear reconstruction problem, where you will find that the sensitivity matrix A is very well conditioned, with the A’A matrix being close to diagonal. The DAS beamforming algorithm is *wrong* in the sense that it effectively multiplies the transpose matrix by the signal, rather than using the pseudo-inverse; but there’s little juice to squeeze in terms of reconstruction accuracy by this more correct algorithm. All in all, the reconstruction problem is straightforward, although it can be a bit subtle, as the reconstruction takes place in complex space; refer to the notebooks for details.\n\n## Isolating blood flow\n\nOnce we have a big stack of rapidly acquired images, we can use a moving window estimate of the standard deviation to measure how the fine-grained texture of the images changes, a proxy for blood volume. This is known as the Power Doppler (PD) estimate. The name is misleading: PD doesn’t measure a Doppler frequency shift, only the change in magnitude and phase of the images. It is sensitive to motion along the beam *and* orthogonal to it. It is calculated by taking the standard deviation of the complex reconstruction. Simulations and analytical solutions show that PD is mostly sensitive to the number of moving scatterers, and less so to their speed (see the notebooks for subtleties).\n\nOne problem with the Power Doppler estimate is that it is quite sensitive to global movement that is not due to the hemodynamic response (e.g., holding the camera in the hand, in the Brooklyn Bridge analogy). In brain imaging, everything moves: pulse and respiration move the probe and the tissue (remember, brains have the texture of warm butter!), and the movement of the subject is communicated through the head. The classic solution is called clutter filtering: it takes raw image sequences and uses the SVD to find the leading terms in a decomposition across time and space, then discards them.\n\nIn practice, clutter filtering frequently discards >99% of the variance in the image sequence. It’s quite a blunt instrument, and it probably discards far more than it needs to. A better solution is to directly estimate and cancel the movement using a piecewise rigid motion estimator; for example, using [NoRMCorre](https://github.com/flatironinstitute/NoRMCorre?tab=readme-ov-file), otherwise used for calcium imaging analysis, to correct for global motion before a (less harsh) clutter filter.\n\nIt’s also not clear that the standard deviation is the best estimate of local red blood cell movement. We can take inspiration from other disciplines here: in diffuse correlation spectroscopy, which uses the destructive and constructive interference of light to infer something about the sample, the movement of scatterers is read off from the *autocorrelation* function of the signal, not from its standard deviation. If the Power Doppler signal is mostly sensitive to the number of scatterers—cerebral blood volume—the autocorrelation decay time is more proportional to the velocity—cerebral blood flow. This could give an orthogonal channel to study, as demonstrated by [Tang et al. (2020)](https://pmc.ncbi.nlm.nih.gov/articles/PMC7509671/).\n\nFinally, one could imagine forgoing all this signal processing and just using an end-to-end deep learning method to embed image sequences for decoding. Unfortunately, there is very little of this data freely available, precluding one from training a foundation model for fUSI on real data, but simulations could do the trick: you just have to make sure that the deep learning method can ultimately process thousands of frames per second.\n\n## Analyzing the data\n\nSuppose, as in Rabut et al. (2024), that we measure the decluttered Power Doppler signal over motor cortex in a paradigm where a patient is instructed to use a joystick to draw an image in different ways. This kind of time series, where the activity is expected to be time-locked to an external signal, can be fruitfully analyzed with some of the same methods as one would use to analyze fMRI. That means GLMs, decoding and encoding models, MVPA, and all their variants. We can use a GLM, for example, to find voxels which are modulated by movement intent; we can use a decoder (e.g. logistic regression) to determine in which direction the patient attempted to move a joystick.\n\nOne of the interesting open questions is how much the extra resolution of fUSI buys you compared to fMRI in terms of decoding accuracy. It probably depends on the variable to be decoded. Some important variables are smoothly encoded on the cortical manifold, for example, retinotopy and somatotopy. Other aspects are interspersed in a far more granular way; orientation selectivity in primary visual cortex is one example. [It’s natural to think that more granularly encoded variables will benefit more than those that are smoothly encoded](https://www.jneurosci.org/content/31/13/4792). Answering this question is important to bound the applicability of fUSI for brain-machine interfaces—what’s the granularity of the movement that you can decode?—or to neuropsychiatric disorders—can we decode subjective value from a heterogeneous population of neurons to nudge it?\n\nBeyond a certain resolution, however, the signal becomes limited by the topology of vasculature. The signal at the level of the capillaries, which feed and drain a handful of neurons, is maximally selective, but it is tiny. By contrast, big veins carry signals large enough to be picked up by fUSI, but because they average over a large drainage area with mixed selectivity, they can be poorly selective. It’s a little like reading the level of pollutants from the Mississippi River using remote sensing by satellite: you will get a lot of signal in New Orleans, where the river is a mile wide, but it wouldn’t resolve the source of the pollutants, because the watershed at that point covers a good chunk of the continent.\n\nOdds are that venules are at the sweet spot. Indeed, GLM-derived activation images of fUSI show striking patterns: it seems the best modulated signals are not on top of (clearly visible) large veins, but rather they are in the periphery of these veins, which I presume are venules.\n\n# Conclusion\n\nIt’s early days for fUSI. The modality has a high ceiling, with a lot of room to grow. Even without access to special equipment, there are several projects, which I’ve alluded to here, that a budding computational neuroscientist or machine learning engineer could do to improve its analysis.\n\nA big part of the growth path of fUSI is to make it more accessible. When I first started looking at fUSI about a year ago, AI agents couldn’t properly simulate ultrasound propagation or correctly code delay-and-sum. That’s why I worked on a series of four notebooks to showcase what I discuss in this essay and convince myself I understand how it works: ultrasound propagation, structural reconstruction, functional reconstruction, and real data analysis. [The notebooks are open source; your contributions and PRs are always welcome](https://github.com/patrickmineault/fus).\n\n[1](#footnote-anchor-1)\n\nThe Midjourney scanner is another example of a structural ultrasound, but unlike a sonogram, it uses attenuation and delay in a transmission configuration in addition to backscatter.\n\n[2](#footnote-anchor-2)\n\nProvided the rigid skull is out of the way.\n\n[3](#footnote-anchor-3)\n\nIt’s a fairly precise analogy, except that the source of contrast is the random relative angle of ambient light and the surface of the water rather than speckle.\n\n[4](#footnote-anchor-4)\n\nA related set of techniques includes video elastography and Eulerian magnification. You can apply these methods to regular images, ultrasound, or optical coherence tomography.", "url": "https://wpnews.pro/news/functional-ultrasound-imaging-from-scratch", "canonical_source": "https://www.neuroai.science/p/functional-ultrasound-imaging-from", "published_at": "2026-09-30 20:06:39+00:00", "updated_at": "2026-09-30 20:19:38.798769+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "neural-networks", "developer-tools"], "entities": ["Patrick Mineault", "Functional ultrasound imaging (fUSI)", "low-intensity focused ultrasound (LIFU)", "GitHub", "Matlab", "Python"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/functional-ultrasound-imaging-from-scratch", "markdown": "https://wpnews.pro/news/functional-ultrasound-imaging-from-scratch.md", "text": "https://wpnews.pro/news/functional-ultrasound-imaging-from-scratch.txt", "jsonld": "https://wpnews.pro/news/functional-ultrasound-imaging-from-scratch.jsonld"}}