Transition and turbulence

The best estimate of a scale assumes its shape

There are four common ways to read an integral time scale off a turbulence record, and on the right signal the best of them is four times more precise than the usual one. On a signal whose correlation has a different shape it is off by a factor that no length of record reveals — and one of them turns out to be measuring the sampling rate rather than the flow.

Worth reading first: A record is as long as its integral scales · The moment a spectrum cannot hold.

A record is as long as its integral scales makes the integral time scale the unit of every error in a turbulence measurement, and then finds it the hardest statistic in the record to measure. The usual estimate integrates the sample autocorrelation to its first zero crossing, because the integral over all lags is exactly zero for every record; it is biased low on short records and uncertain by ten per cent after three thousand integral scales. The essay ends by naming two alternatives, a model fitted to the autocorrelation and the scatter of block means, and asking which is best.

The answer turns out to depend less on the estimators than on what each one assumes, and the assumptions are the interesting part. An estimator that assumes the correlation’s shape is far more precise than one that does not — when the shape is right. When it is wrong, the precision is still there and the answer is wrong, and nothing in the record says so.

Four estimators and what each takes on trust

The first zero integrates the sample autocorrelation from zero lag to its first crossing of zero. It uses the definition of the integral scale and assumes nothing about the correlation’s shape, only that the correlation stays positive until it has nearly finished.

The exponential fit assumes ρ(τ)=e−τ/TI\rho(\tau) = e^{-\tau/T_I} and fits the slope of ln⁡ρ^\ln\hat\rho against lag, by least squares through the origin, over the lags where the sample correlation is above 0.3.

The lag-one estimate assumes the same shape and uses only the first lag: TI=−Δt/ln⁡ρ^(Δt)T_I = -\Delta t/\ln\hat\rho(\Delta t). For an exponentially correlated process this is essentially the likelihood estimate, and it should be very good.

The block estimate cuts the record into mm blocks of length BB and asks what integral scale would give the block means the scatter they have: since the variance of a mean over BB is 2σ2TI/B2\sigma^2 T_I/B for long blocks, TI=Bsb2/2σ^2T_I = B s_b^2/2\hat\sigma^2. It assumes nothing about the shape either. It uses the integral scale’s one job — setting the error of a mean — to measure it.

Two signals with one integral scale

Each estimator was run on three hundred records at each of five lengths from thirty to three thousand integral scales, five samples per integral scale, and on two different processes.

Two correlations with the same integral scale. The autocorrelation of the exponential process and of the smooth one, against lag in integral scales, with the smooth process's correlation measured from one long generated record. Both integrate to one. The exponential falls with a corner at zero lag; the smooth one starts flat, as the correlation of any real velocity does, and then falls faster.
Fig. 1 The autocorrelation of the exponential process and of the smooth one against lag, the smooth one also measured from a long generated record. Both integrate to one. The exponential falls with a corner at zero lag; the smooth one starts flat, as any real velocity’s correlation does, and falls faster later.

The first is the exponentially correlated process of the essay before. The second is the same process passed through a second, identical first-order filter, which gives a correlation (1+∣τ∣/τ0)e−∣τ∣/τ0(1 + |\tau|/\tau_0)e^{-|\tau|/\tau_0} with τ0=12\tau_0 = \tfrac12: the same integral scale, one, and a correlation that starts flat at zero lag instead of falling with a corner. The flat start is not a curiosity. A velocity that has derivatives has a correlation whose slope at zero lag is zero, and the curvature there is the Taylor microscale — the scale of the velocity gradients, set by viscosity. Every real turbulent velocity has one. The exponential, with its corner, describes a signal with no derivative anywhere, which no flow produces.

The generated smooth record matches its intended correlation to 1.1×10−31.1 \times 10^{-3} at four lags and has a variance of 1.004.

On its own model, an assumption is precision

Precision bought with an assumption. The root-mean-square error of each estimate, bias and scatter together, as a fraction of the true integral scale, against record length — dashed on the exponential process, solid on the smooth one. The lag-one and exponential-fit estimates are the most precise on the process whose shape they assume and stop improving on the other; the first zero is shape-free and slow; twenty blocks never beat 30 per cent.
Fig. 2 The root-mean-square error of each estimate, bias and scatter together, against record length — dashed on the exponential process, solid on the smooth one. The shape-assuming estimates are the most precise on the process whose shape they assume and stop improving on the other.

On the exponential process the ranking is what the theory of estimation predicts. The lag-one estimate is the best: its error is 9.0 per cent at three hundred integral scales, 4.9 at a thousand and 2.7 at three thousand. The exponential fit, which uses a few more lags and so a little more noise from the correlation’s tail, is close behind at 11.6, 6.5 and 3.7. The first zero, which integrates through the noisy tail as well, manages 27, 15 and 10 per cent. Knowing the shape is worth a factor of three to four in precision, the equivalent of a record ten times longer.

The block estimate with twenty blocks is flat at about thirty per cent at every length. Twenty block means give a variance known only to 2/19\sqrt{2/19} of itself, about a third, however long the blocks are. That is a defect of fixing the number of blocks, which a later section repairs.

Off its model, the assumption is a bias the record cannot show

What each estimate reports on average, when the correlation is smooth. The average over three hundred records of three estimates of the integral scale of the smooth process, as a fraction of the truth, against record length: the first zero of the sample autocorrelation, an exponential fit to it, and twenty block means. The first zero and the blocks converge on one; the exponential fit settles at about 1.18 however long the record, because the shape it assumes is wrong.
Fig. 3 The average of three estimates of the smooth process’s integral scale against record length. The first zero and the blocks converge on the truth; the exponential fit settles at about 1.18 however long the record, because the shape it assumes is wrong.

On the smooth process the exponential fit is just as precise as before — its scatter at three thousand integral scales is 3.3 per cent — and it is precisely wrong. It settles at 1.18 of the true integral scale and stays there as the record lengthens: 1.157 at a hundred scales, 1.178 at three hundred, 1.186 at a thousand, 1.178 at three thousand. Its error of eighteen per cent is not noise, and so a longer record does not reduce it; the record would have to be examined for the shape of its correlation, not for its length, to find it.

The shape-free estimators behave differently. The first zero converges on the truth from slightly above, 1.089 at three hundred scales and 1.019 at three thousand. The twenty-block estimate converges from below. Neither cares which process it is given: at a thousand integral scales the twenty-block means average 0.964 on one process and 0.987 on the other, a difference within their scatter, while the exponential fit separates by 0.19.

Where the eighteen per cent comes from

The fit’s bias has nothing to do with the records. Apply the same fit to the exact correlations, with no noise at all, and it returns the same number.

An exponential fit returns whatever its window makes it. The integral scale an exponential fit reports when it is applied to the exact correlation — no record, no noise — against the correlation level at which the fit is cut off. On the exponential correlation every window returns one. On the smooth correlation, a fit to the flat top reports more than twice the truth and a fit that runs well into the tail reports less than it; the fit crosses the truth at a window nobody would choose on purpose.
Fig. 4 The scale an exponential fit reports when applied to the exact correlation, against the lowest correlation level included in the fit. On the exponential every window returns one. On the smooth correlation a fit to the flat top reports more than twice the truth and a fit well into the tail reports less.

An exponential fitted to a curve that starts flat has to choose between matching the top and matching the tail, and the window decides which. Fitted only where the correlation exceeds 0.8 — the flat top — it reports 2.29 times the true scale, because a gentle initial fall looks like a long decay. Fitted down to 0.5 it reports 1.44; to 0.3, the window used above, 1.181, which is the 1.18 the records converged on; to 0.1, 0.958; to 0.05, 0.888. On the exponential correlation every one of these windows returns exactly one.

So the bias is not a property of the estimator alone but of the estimator, the window and the shape together, and the window is a choice the analyst makes. Two laboratories fitting exponentials to the same smooth correlation, one down to half its peak and one down to a tenth, would report integral scales fifty per cent apart and each would be precise. The only protection is the one the shape-free estimators provide: an answer that does not move when the window does.

What the choice is worth in record

Precision translates into record length, and record length into hours, which is where the temptation to assume a shape comes from. To know the integral scale to ten per cent, the first zero needs about three thousand integral scales of record; the lag-one estimate on a genuinely exponential signal needs about three hundred, and the exponential fit about five hundred. In a wind tunnel, where the integral time is ten milliseconds, that is the difference between thirty seconds of record and five, and nobody minds. In the atmospheric surface layer, where it is about a minute, it is the difference between two days of stationary weather, which never happens, and five hours, which occasionally does.

That is why atmospheric integral scales are usually fitted rather than integrated, and why they carry a bias of the kind found here that no single record can reveal. The fit is not a mistake; it is the only way to get a number at all from records that short. What it needs is the sentence saying which shape and which window were assumed, so that the number can be moved when the assumption is.

The estimate that measures the sampling rate

The lag-one estimate fails more interestingly than the fit, because on a smooth signal it is not measuring the integral scale at all.

On a smooth signal, sampling faster makes the lag-one scale longer. The average estimate of the integral scale from records three hundred scales long of the smooth process, against the number of samples per integral scale. The first zero and the exponential fit do not care how fast the record was sampled. The lag-one estimate grows in proportion to the sampling rate, because it is reading the curvature of the correlation at zero lag — the microscale — and dividing by the sample spacing.
Fig. 5 The average estimate from records three hundred scales long of the smooth process, against the number of samples per integral scale. The first zero and the exponential fit do not care how fast the record was sampled. The lag-one estimate grows in proportion to the sampling rate.

For a smooth signal the correlation near zero lag is 1−τ2/2λ21 - \tau^2/2\lambda^2, where λ\lambda is the microscale. Then −ln⁡ρ^(Δt)≈Δt2/2λ2-\ln\hat\rho(\Delta t) \approx \Delta t^2/2\lambda^2, and the lag-one estimate is 2λ2/Δt2\lambda^2/\Delta t: the square of the microscale divided by the sample spacing. It has nothing to do with the integral scale. The smooth process has λ=0.5\lambda = 0.5, so the estimate should be 0.5/Δt0.5/\Delta t once Δt\Delta t is small, and it is: 3.08 at five samples per integral scale, 5.47 at ten, 10.1 at twenty. Sample twice as fast and the flow’s “integral scale” doubles.

That makes the lag-one estimate unusable on turbulence data, where the sampling is chosen to resolve the smallest scales and so is far faster than the microscale requires: a hot-wire record at a hundred kilohertz in a flow whose integral time is ten milliseconds takes a thousand samples per integral scale. It also explains why the estimate is attractive in other fields, where it works. A signal that genuinely is exponentially correlated at the sampling interval — a count rate, a queue, a random walk in a potential — has no microscale, and there the lag-one estimate is the best there is.

Blocks, and the circle again

The block estimate’s thirty per cent was a choice, not a limit. Its error has two parts that pull against each other through the block length BB. Short blocks have means that are correlated with their neighbours, and the formula 2σ2TI/B2\sigma^2T_I/B is only the long-block limit, so short blocks report too small a scale. Long blocks are few, and the variance of a few means is itself uncertain.

Blocks about ten integral scales long. The root-mean-square error of the block estimate against the length of its blocks, for records of 300, 1000 and 3000 integral scales of the smooth process. Short blocks have correlated means and report too small a scale; long blocks are too few to give a steady variance. The best block length is about ten integral scales at every record length — which needs the answer to choose it.
Fig. 6 The block estimate’s error against its block length, for records of 300, 1000 and 3000 integral scales. Short blocks report too small a scale; long ones are too few. The best block length is about ten integral scales at every record length.

The best compromise is a block length of about ten integral scales, whatever the record’s length. At that length the block estimate’s error is 18 per cent on three hundred scales of record, 13 on a thousand and 10 on three thousand — as good as the first zero, or slightly better on the shorter records, and like it free of any assumption about the shape. At a thousand scales of record, blocks of fifty, ten and two integral scales give errors of 29, 15 and 43 per cent.

The recommendation contains its own difficulty: choosing blocks ten integral scales long requires knowing the integral scale. The circle of the essay before has not been broken, only moved. The way round it in practice is the plateau reading that essay described — plot the block estimate against the block length, and read it where it stops rising — which is a single choice made by looking, and which the shape-free estimators need and the shape-assuming ones do not.

What the integral-scale comparison was checked against. The numbers quoted and their checks: the smooth process against its intended correlation and variance, each estimator on the process it assumes and on the one it does not, and the block estimator's trade between bias and scatter.
Fig. 7 The numbers quoted and their checks: the smooth process against its intended correlation, each estimator on the process it assumes and on the other, and the block estimate’s trade between bias and scatter.

The same number, read off the spectrum

There is a fifth route, and it turns out to be the fourth in disguise. The integral scale is the spectrum at zero frequency divided by the variance — TI=πE(0)/σ2T_I = \pi E(0)/\sigma^2 for a two-sided spectrum normalised so — and a spectrum is something every turbulence record is analysed for anyway. Reading E(0)E(0) off the lowest frequencies of a measured spectrum looks shape-free and cheap.

It is shape-free, and it is exactly as expensive as the blocks. A raw periodogram’s value at any one frequency is uncertain by a hundred per cent of itself however long the record, because each bin is the sum of the squares of two Gaussian numbers. To get a usable value near zero frequency the periodogram must be averaged — over neighbouring frequencies, or over segments of the record, which is Welch’s method — and averaging over segments of length BB is the block estimate by another name: the variance of a segment’s mean is the spectrum integrated over a band of width 1/B1/B around zero. The trade is the same, too. Short segments give many averages of a spectrum smeared over too wide a band; long segments give a sharp band and few averages. The optimum is again a segment several integral scales long, and the achievable precision the same ten to twenty per cent.

That equivalence is useful because it says the choice between shape-free methods is a matter of convenience and not of accuracy. The first zero, the block plateau and the averaged low-frequency spectrum all converge at the same slow rate, all need a judgement made by looking, and all are free of the one error that the fitted methods cannot see. A spectrum and the loads it implies inherit whichever one was used.

What a measurement should report

Three practical rules fall out of the comparison, and each is a statement about assumptions rather than about arithmetic.

Report the estimator. Four defensible methods applied to one smooth record of three hundred integral scales give answers from 0.97 to 3.1 times the truth. An integral scale quoted without its method is not a number that can be compared with another laboratory’s.

Use a shape-free estimate first, and check the shape second. The first zero and the block plateau give an answer good to ten or twenty per cent that no assumption can spoil. If the sample correlation then turns out to be exponential over the lags that matter, a fitted estimate can be trusted for a factor of three in precision; if it is flat at the origin, it cannot.

Never read a turbulence integral scale off the first lag. It is the microscale squared over the sample spacing, and a faster anemometer will report a longer scale.

What the picture cannot show

Gaussian records. Both processes are Gaussian. A turbulent velocity is close to Gaussian and its correlation’s shape, which is what matters here, is the property carried over; its derivatives are not Gaussian, which matters to microscale estimates and not to these.

Two shapes. A real correlation is neither of these: it is flat at the origin over the microscale, roughly exponential over the middle, and sometimes negative at large lags, as a record from a flow with a dominant eddy frequency is. An exponential fit’s bias on such a record is of the size found here and of unknown sign.

A stationary flow. Every record is statistically steady. A drifting flow adds a long-lag correlation that the first zero and the blocks both mistake for a longer integral scale, which is the non-stationarity the essay before warned about and which no estimator here detects.

The convention the numbers depend on

The true integral scale of both processes is one, and every estimate is a fraction of it. Records are sampled at five samples per integral scale except where the sampling rate is the variable. “Error” is the root-mean-square difference from the truth across records, bias and scatter together. The exponential fit uses lags where the sample correlation exceeds 0.3; the block estimate uses the record’s own mean and variance.

Who found it, and when

The first-zero convention is as old as the integral scale, which is G. I. Taylor’s from the 1920s. Estimating an error from the scatter of block means is Bartlett’s method of 1946 for spectra, and its use to measure a correlation time from the plateau of the block variance was popularised in molecular simulation by Flyvbjerg and Petersen in 1989, where records are also long, correlated and expensive. Comparisons of integral-scale estimators on turbulence data have repeatedly found the spread between methods to exceed the scatter of any one, which is the empirical form of the result here.

It sits beside the moment a spectrum cannot hold, where a correlation’s shape at the origin decides whether a spectral moment exists, and the range a real Reynolds number does not have, where a model spectrum’s shape is the assumption doing the work.

Still open: a field instead of a record

Every estimate here reads a time series at one point. A particle-image velocimetry frame or a simulation snapshot gives a field instead, and the integral scale is then a length estimated from spatial correlations across a finite window — the window every vector is averaged over sets the smallest separation and the frame’s size the largest. The calculation that follows repeats the comparison in two dimensions: fields of known integral length generated exactly, the four estimators applied across windows of increasing size, and the question whether a snapshot of given size holds more or fewer independent values than a record of the same number of samples. The answer decides how many snapshots the exact laws’ constants need before their error bars are honest.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

AveragingConvergenceCorrelationIntegral scaleMeasurementModel validityProbability distributionSamplingSpectrumStatistics