The randomness that is not in the equations
Worth reading first: What averaging costs · The number that is not a number.
Open any book on turbulence and the vocabulary is statistical from the first page: mean velocities, fluctuations, variances, correlations, spectra, probability density functions. Open the equations the book is about and there is nothing statistical in them at all. The Navier–Stokes equations are as deterministic as the equations of a pendulum: given the fluid, the boundaries and the state at one instant, the state at every later instant follows.
Something has to reconcile those two facts, and the reconciliation is not that a random term was left out.
What the equations actually contain
The momentum equation for a Newtonian fluid is
with a body force such as gravity. There is no term with a random variable in it, no noise, no probability. A calculation started from given data reproduces itself exactly on a second run, which every practitioner of direct numerical simulation relies on daily; the same code with the same data on the same machine gives the same answer to the last bit.
Nor is the equation ill-posed in a way that would let several futures follow from one present. Whether solutions remain smooth for all time in three dimensions is famously unsettled, and that is a question about the existence of a solution rather than about there being several — nothing in the open problem suggests a random element.
So the statistics did not come from the physics. They came from somewhere else, and there are three candidates, of which two are real.
Where they do come from
Sensitivity to initial data. A turbulent flow amplifies small differences exponentially. Two states differing by a millionth are indistinguishable now and completely different in a few large-eddy turnover times, and since no measurement of an initial condition is exact, the actual trajectory a real flow takes is not predictable beyond a horizon. It is determined and unknown, which is a different thing from undetermined.
Ignorance of the boundary and forcing data. No two runs of a laboratory experiment have the same inflow, the same wall roughness at the micron scale, or the same vibration. An experimenter repeating a measurement is not repeating a solution; they are drawing a new member of a family of solutions whose data differ in ways nobody has recorded.
That second point has a consequence for experiments that is easy to underrate. A wind tunnel’s turbulence level is a specification precisely because it is part of the initial data: the same model in two tunnels transitions at different Reynolds numbers, and the number quoted for transition is really a property of the disturbance environment rather than of the fluid.
And the sheer number of degrees of freedom. Even if the data were known exactly, a turbulent flow at a laboratory Reynolds number has independent modes — a number this collection computes and which reaches for a wing. A description that reports every one of them is not a description.
The first two are why turbulence is unpredictable; the third is why it is described statistically even where it is predictable. They are separate arguments and they are often run together.
What an ensemble average is an average over
The angle brackets in every equation of turbulence theory denote an ensemble average, and it is worth being exact about what the ensemble contains. It is a set of realisations with the same boundary conditions and different initial data — or, equivalently, the same experiment repeated on different days.
That is not an average over noise, because there is no noise. It is an average over ignorance, and it has a property that matters: the ensemble is a construct of the analyst, and different choices of what to hold fixed give different averages. A phase average over the cycle of a rotating machine, a conditional average over the turbulent state at an intermittent edge, and a plain time average are three different ensembles of the same flow, and their means are three different fields.
The assumption that lets one long record stand for an ensemble
The section above defines the average over an ensemble of initial data, and nobody has such an ensemble. What every experiment and every simulation actually does is average over time, along a single trajectory, and substituting the one for the other is an assumption with a name and with failure modes worth knowing.
The assumption is ergodicity: that a single long trajectory visits the accessible states with the same frequencies the ensemble assigns to them, so that a time average converges to the ensemble average. Where it holds, one record is as good as many, and the whole apparatus of turbulence measurement rests on it. It is almost never stated and almost never tested.
Two conditions have to hold for it to be reasonable, and both are checkable.
The flow must be stationary. Its statistics must not depend on when the record was taken, which rules out decaying turbulence outright — a grid-generated field whose energy is falling has different statistics at every instant, so its ensemble must be built from repeated runs rather than from a long record, and every measurement of decaying turbulence is made that way for this reason.
And the record must be long compared with the slowest correlation in it. That is the harder condition, because the slowest correlation is the one least likely to be noticed. A record covering thousands of large-eddy turnovers looks abundant and is still short if the flow has a mode with a period of minutes — a slowly oscillating separation bubble, a switching of a wake between two asymmetric states, a drift in the facility’s temperature. The statistic converges beautifully and converges to the wrong thing, and its own error bar, computed from the fluctuations it can see, says nothing about the mode it cannot.
The bistable case is the sharpest and it connects directly to arithmetic this collection has already done. A flow that spends long intervals in one of two states — the wake of a blunt body that switches between two mirror-image configurations is the standard example — has a time-averaged mean that is the weighted average of the two, and neither state resembles it. The mean is symmetric and the flow is never symmetric. That is exactly the two-state arithmetic an intermittent edge produces, arriving in a flow with no interface in it, and it produces the same artefacts: a variance inflated by the switching, and a “mean field” that is a description of nothing.
So the essay’s remark that the ensemble is a construct of the analyst has a sharper form. The ensemble is a construct, the time average is a different construct, and the identification of the two is an unstated hypothesis about the slowest thing in the flow. Testing it costs one calculation — split the record in halves and compare — and the fact that this is rarely done is the reason so many disagreements between facilities turn out, on inspection, to be disagreements about how long anybody watched.
The point where the misconception does damage
If turbulence were noisy, two things would follow that do not.
A model would be a filter rather than an approximation. Adding random forcing to a calculation would then be more faithful than leaving it out, and stochastic models would be the natural class. They are used, and they are used as models of ignorance rather than as models of the fluid — the distinction shows up the moment a result depends on the noise’s amplitude, which is a modelling choice with no counterpart in the equations.
And repeatability would be evidence of something being suppressed. In fact a direct simulation repeats exactly, and its statistics converge to the same values as a laboratory experiment whose individual realisations are completely different. That agreement is the strongest available evidence that the statistics are properties of the equations rather than of any noise process, and it is routinely used to validate codes.
It also changes what a measurement is for. If the flow were noisy, a long record would be a sample from a distribution the fluid possesses; since it is not, a long record is a sample from a distribution the experiment possesses — and two facilities running the same nominal experiment sample different distributions. That is exactly the difficulty intermittency measurements meet when their thresholds differ between laboratories, and it is why so much of this subject’s experimental literature is careful about facilities in a way that other fields are not.
There is a third consequence that is subtler and more useful. Because the equations are deterministic, a statistic is not free to be anything: it inherits constraints from the dynamics. The four-fifths law is exactly such a constraint — a statement about a third moment that follows from the equations of motion — and no stochastic model produces it by accident. A description that treated turbulence as noise would have no reason to expect any exact statistical law at all.
How long a flow can be predicted, and what sets it
The horizon is worth quantifying, because it is the practical content of everything above. If two states differ by and separate as , then the time until the difference reaches the size of the flow itself is
and the logarithm is the whole difficulty: improving the initial data by a factor of a thousand buys only seven more -foldings of prediction. In the atmosphere is a day or two and is of order ten, which is where the fortnight-long limit on weather forecasting comes from — a limit that better instruments cannot lift, only slide.
For a laboratory flow the growth rate is set by the large-eddy turnover time, so the horizon is a few turnovers whatever the Reynolds number. That is why a direct simulation is compared with an experiment through statistics rather than instant by instant: the two calculations agree about every mean and about no individual eddy, and both are right.
Chaos is not the same claim as turbulence
The word chaos does a great deal of work in popular accounts and it is worth separating two statements that both use it.
Sensitive dependence is a property of a trajectory. It says two nearby states separate exponentially, at a rate measured by a Lyapunov exponent, and it applies to systems with three degrees of freedom as readily as to a fluid.
Turbulence is a property of a flow with very many degrees of freedom, whose spectrum spans decades of scale, and a cascade of energy between them. A chaotic system with three variables has no cascade, no inertial range and no dissipation that survives the vanishing of viscosity.
The Lorenz system is chaotic and is not turbulent. A turbulent flow is chaotic and has the structure the rest of this field is about. Conflating them was a fashionable error in the 1980s and the correction is now standard.
The reverse case, which makes the point cleanly
A flow can also be deterministic, well mixed and not chaotic at all. The blinking vortex — two stirring rods used alternately — mixes a blob of dye into a folded ribbon whose spectrum broadens steadily, and it does so with no randomness of any kind.
It is the same demonstration this collection makes with a creeping flow that unmixes itself, where reversibility is a property of the equations rather than of the stirring, and where a drop of dye smeared through a whole annulus reassembles into a drop.
That example separates the ideas cleanly. Mixing does not require randomness; complexity does not require randomness; a broad spectrum does not require randomness. What randomness would supply — and what a deterministic flow supplies instead through sensitivity — is only the unpredictability.
The one place a random term is honest
There is a setting in which adding noise to the equations is not a confusion, and separating it from the misconception is worth a paragraph.
A model that has discarded scales may honestly represent them stochastically. A large-eddy simulation resolves the large scales and models the small ones; the small ones are not known, their effect on the resolved field is genuinely uncertain, and a stochastic sub-grid model is a statement about that uncertainty rather than about the fluid. The same is true of a Langevin model for the motion of a parcel in a dispersion calculation: the parcel’s path is deterministic and unknown, and the model’s noise represents the not-knowing.
The distinction is the one between a fluid that is random and a description that is incomplete. Every honest stochastic model in this subject is of the second kind, and each of them carries a parameter — the noise amplitude — that has to be calibrated, which is the tell: a real physical noise would have a strength set by the physics, as Brownian motion’s is set by the temperature.
Brownian motion is the contrast that settles it. There the randomness is real and its amplitude is fixed by , through a fluctuation–dissipation relation that connects it to the viscosity. Nothing of the kind exists for turbulence, and looking for one is a good way to see why the analogy fails.
What the picture cannot show
The Lorenz system is not a fluid. It is a three-mode truncation of a convection problem, and this collection labels it as one wherever it is drawn. Its attractor is a beautiful object and it is not a flow; the properties it demonstrates — sensitivity, a bounded attractor, a positive Lyapunov exponent — are properties of that system, and the claim made here is only that a deterministic system can have them.
No figure on this site contains a turbulent field, so nothing here demonstrates that Navier–Stokes turbulence is chaotic. That it is has been established by direct simulation, with measured Lyapunov exponents and a predictability horizon of a few large-eddy times, and it is imported as a fact rather than computed.
The predictability numbers are imported. The Lyapunov exponent of a turbulent flow, and hence its horizon, comes from published direct simulations; nothing here measures one for a fluid. What is computed is the exponent of a three-variable system, which establishes that a deterministic system can have a positive one and nothing more.
And the open mathematical question is genuinely open. Whether three-dimensional solutions remain smooth for all time is unsettled, and a proof of blow-up would mean the equations do not determine the flow past a certain moment. That would be a much more interesting development than noise, and it would not make turbulence random either.
Who noticed, and when
Reynolds’ 1883 experiments established that a flow becomes irregular past a critical value of his number, and Reynolds himself framed the resulting description statistically in 1895, which is where the decomposition and the stresses come from. The identification of that irregularity with deterministic sensitive dependence took until Lorenz in 1963 and the experiments and analysis of the 1970s; before that the dominant picture was Landau’s, in which turbulence is a superposition of many incommensurate frequencies acquired one at a time — a picture that is deterministic too, and that turned out to be wrong for a different reason.
The surprising connection is with the other essays in this collection about averages. Each is about an average that does not commute with the physics, and this one is about where the average came from in the first place. Turbulence is described statistically because nobody can write down the initial data, not because the fluid is throwing dice — and that matters for what a statistic is entitled to say: it is a statement about an ensemble the analyst constructed, and its properties are constrained by equations that know nothing about ensembles.
Where the ladder goes next
This is the misconception’s own rung, and the ladder it sits beside is the transition one, where the same question is asked quantitatively: what a critical Reynolds number is a number for, given that a pipe stays laminar to a hundred thousand if it is quiet enough. Above it lies predictability itself — how far ahead a flow can be forecast, and why that horizon is a property of the flow rather than of the computer.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- How far a parcel gets — both name closure, measurement, mixing, statistics, turbulence
- Two averages of one flow — both name averaging, ensemble, measurement, turbulence
- What a mean profile cannot tell anybody — both name averaging, closure, measurement, turbulence
- A closure with no memory at all — both name closure, measurement, turbulence
- A dissipation that lags its production — both name closure, measurement, turbulence
- A scalar is a record of where its fluid was — both name measurement, mixing, predictability
Named objects
A dashed tag is an object no other essay names yet.
AveragingChaosClosureDeterminismEnsembleLorenzMeasurementMixingPredictabilitySensitivityStatisticsTurbulence