The best estimate of a scale assumes its shape
Worth reading first: A record is as long as its integral scales · The moment a spectrum cannot hold.
A record is as long as its integral scales makes the integral time scale the unit of every error in a turbulence measurement, and then finds it the hardest statistic in the record to measure. The usual estimate integrates the sample autocorrelation to its first zero crossing, because the integral over all lags is exactly zero for every record; it is biased low on short records and uncertain by ten per cent after three thousand integral scales. The essay ends by naming two alternatives, a model fitted to the autocorrelation and the scatter of block means, and asking which is best.
The answer turns out to depend less on the estimators than on what each one assumes, and the assumptions are the interesting part. An estimator that assumes the correlation’s shape is far more precise than one that does not — when the shape is right. When it is wrong, the precision is still there and the answer is wrong, and nothing in the record says so.
Four estimators and what each takes on trust
The first zero integrates the sample autocorrelation from zero lag to its first crossing of zero. It uses the definition of the integral scale and assumes nothing about the correlation’s shape, only that the correlation stays positive until it has nearly finished.
The exponential fit assumes and fits the slope of against lag, by least squares through the origin, over the lags where the sample correlation is above 0.3.
The lag-one estimate assumes the same shape and uses only the first lag: . For an exponentially correlated process this is essentially the likelihood estimate, and it should be very good.
The block estimate cuts the record into blocks of length and asks what integral scale would give the block means the scatter they have: since the variance of a mean over is for long blocks, . It assumes nothing about the shape either. It uses the integral scale’s one job — setting the error of a mean — to measure it.
Two signals with one integral scale
Each estimator was run on three hundred records at each of five lengths from thirty to three thousand integral scales, five samples per integral scale, and on two different processes.
The first is the exponentially correlated process of the essay before. The second is the same process passed through a second, identical first-order filter, which gives a correlation with : the same integral scale, one, and a correlation that starts flat at zero lag instead of falling with a corner. The flat start is not a curiosity. A velocity that has derivatives has a correlation whose slope at zero lag is zero, and the curvature there is the Taylor microscale — the scale of the velocity gradients, set by viscosity. Every real turbulent velocity has one. The exponential, with its corner, describes a signal with no derivative anywhere, which no flow produces.
The generated smooth record matches its intended correlation to at four lags and has a variance of 1.004.
On its own model, an assumption is precision
On the exponential process the ranking is what the theory of estimation predicts. The lag-one estimate is the best: its error is 9.0 per cent at three hundred integral scales, 4.9 at a thousand and 2.7 at three thousand. The exponential fit, which uses a few more lags and so a little more noise from the correlation’s tail, is close behind at 11.6, 6.5 and 3.7. The first zero, which integrates through the noisy tail as well, manages 27, 15 and 10 per cent. Knowing the shape is worth a factor of three to four in precision, the equivalent of a record ten times longer.
The block estimate with twenty blocks is flat at about thirty per cent at every length. Twenty block means give a variance known only to of itself, about a third, however long the blocks are. That is a defect of fixing the number of blocks, which a later section repairs.
Off its model, the assumption is a bias the record cannot show
On the smooth process the exponential fit is just as precise as before — its scatter at three thousand integral scales is 3.3 per cent — and it is precisely wrong. It settles at 1.18 of the true integral scale and stays there as the record lengthens: 1.157 at a hundred scales, 1.178 at three hundred, 1.186 at a thousand, 1.178 at three thousand. Its error of eighteen per cent is not noise, and so a longer record does not reduce it; the record would have to be examined for the shape of its correlation, not for its length, to find it.
The shape-free estimators behave differently. The first zero converges on the truth from slightly above, 1.089 at three hundred scales and 1.019 at three thousand. The twenty-block estimate converges from below. Neither cares which process it is given: at a thousand integral scales the twenty-block means average 0.964 on one process and 0.987 on the other, a difference within their scatter, while the exponential fit separates by 0.19.
Where the eighteen per cent comes from
The fit’s bias has nothing to do with the records. Apply the same fit to the exact correlations, with no noise at all, and it returns the same number.
An exponential fitted to a curve that starts flat has to choose between matching the top and matching the tail, and the window decides which. Fitted only where the correlation exceeds 0.8 — the flat top — it reports 2.29 times the true scale, because a gentle initial fall looks like a long decay. Fitted down to 0.5 it reports 1.44; to 0.3, the window used above, 1.181, which is the 1.18 the records converged on; to 0.1, 0.958; to 0.05, 0.888. On the exponential correlation every one of these windows returns exactly one.
So the bias is not a property of the estimator alone but of the estimator, the window and the shape together, and the window is a choice the analyst makes. Two laboratories fitting exponentials to the same smooth correlation, one down to half its peak and one down to a tenth, would report integral scales fifty per cent apart and each would be precise. The only protection is the one the shape-free estimators provide: an answer that does not move when the window does.
What the choice is worth in record
Precision translates into record length, and record length into hours, which is where the temptation to assume a shape comes from. To know the integral scale to ten per cent, the first zero needs about three thousand integral scales of record; the lag-one estimate on a genuinely exponential signal needs about three hundred, and the exponential fit about five hundred. In a wind tunnel, where the integral time is ten milliseconds, that is the difference between thirty seconds of record and five, and nobody minds. In the atmospheric surface layer, where it is about a minute, it is the difference between two days of stationary weather, which never happens, and five hours, which occasionally does.
That is why atmospheric integral scales are usually fitted rather than integrated, and why they carry a bias of the kind found here that no single record can reveal. The fit is not a mistake; it is the only way to get a number at all from records that short. What it needs is the sentence saying which shape and which window were assumed, so that the number can be moved when the assumption is.
The estimate that measures the sampling rate
The lag-one estimate fails more interestingly than the fit, because on a smooth signal it is not measuring the integral scale at all.
For a smooth signal the correlation near zero lag is , where is the microscale. Then , and the lag-one estimate is : the square of the microscale divided by the sample spacing. It has nothing to do with the integral scale. The smooth process has , so the estimate should be once is small, and it is: 3.08 at five samples per integral scale, 5.47 at ten, 10.1 at twenty. Sample twice as fast and the flow’s “integral scale” doubles.
That makes the lag-one estimate unusable on turbulence data, where the sampling is chosen to resolve the smallest scales and so is far faster than the microscale requires: a hot-wire record at a hundred kilohertz in a flow whose integral time is ten milliseconds takes a thousand samples per integral scale. It also explains why the estimate is attractive in other fields, where it works. A signal that genuinely is exponentially correlated at the sampling interval — a count rate, a queue, a random walk in a potential — has no microscale, and there the lag-one estimate is the best there is.
Blocks, and the circle again
The block estimate’s thirty per cent was a choice, not a limit. Its error has two parts that pull against each other through the block length . Short blocks have means that are correlated with their neighbours, and the formula is only the long-block limit, so short blocks report too small a scale. Long blocks are few, and the variance of a few means is itself uncertain.
The best compromise is a block length of about ten integral scales, whatever the record’s length. At that length the block estimate’s error is 18 per cent on three hundred scales of record, 13 on a thousand and 10 on three thousand — as good as the first zero, or slightly better on the shorter records, and like it free of any assumption about the shape. At a thousand scales of record, blocks of fifty, ten and two integral scales give errors of 29, 15 and 43 per cent.
The recommendation contains its own difficulty: choosing blocks ten integral scales long requires knowing the integral scale. The circle of the essay before has not been broken, only moved. The way round it in practice is the plateau reading that essay described — plot the block estimate against the block length, and read it where it stops rising — which is a single choice made by looking, and which the shape-free estimators need and the shape-assuming ones do not.
The same number, read off the spectrum
There is a fifth route, and it turns out to be the fourth in disguise. The integral scale is the spectrum at zero frequency divided by the variance — for a two-sided spectrum normalised so — and a spectrum is something every turbulence record is analysed for anyway. Reading off the lowest frequencies of a measured spectrum looks shape-free and cheap.
It is shape-free, and it is exactly as expensive as the blocks. A raw periodogram’s value at any one frequency is uncertain by a hundred per cent of itself however long the record, because each bin is the sum of the squares of two Gaussian numbers. To get a usable value near zero frequency the periodogram must be averaged — over neighbouring frequencies, or over segments of the record, which is Welch’s method — and averaging over segments of length is the block estimate by another name: the variance of a segment’s mean is the spectrum integrated over a band of width around zero. The trade is the same, too. Short segments give many averages of a spectrum smeared over too wide a band; long segments give a sharp band and few averages. The optimum is again a segment several integral scales long, and the achievable precision the same ten to twenty per cent.
That equivalence is useful because it says the choice between shape-free methods is a matter of convenience and not of accuracy. The first zero, the block plateau and the averaged low-frequency spectrum all converge at the same slow rate, all need a judgement made by looking, and all are free of the one error that the fitted methods cannot see. A spectrum and the loads it implies inherit whichever one was used.
What a measurement should report
Three practical rules fall out of the comparison, and each is a statement about assumptions rather than about arithmetic.
Report the estimator. Four defensible methods applied to one smooth record of three hundred integral scales give answers from 0.97 to 3.1 times the truth. An integral scale quoted without its method is not a number that can be compared with another laboratory’s.
Use a shape-free estimate first, and check the shape second. The first zero and the block plateau give an answer good to ten or twenty per cent that no assumption can spoil. If the sample correlation then turns out to be exponential over the lags that matter, a fitted estimate can be trusted for a factor of three in precision; if it is flat at the origin, it cannot.
Never read a turbulence integral scale off the first lag. It is the microscale squared over the sample spacing, and a faster anemometer will report a longer scale.
What the picture cannot show
Gaussian records. Both processes are Gaussian. A turbulent velocity is close to Gaussian and its correlation’s shape, which is what matters here, is the property carried over; its derivatives are not Gaussian, which matters to microscale estimates and not to these.
Two shapes. A real correlation is neither of these: it is flat at the origin over the microscale, roughly exponential over the middle, and sometimes negative at large lags, as a record from a flow with a dominant eddy frequency is. An exponential fit’s bias on such a record is of the size found here and of unknown sign.
A stationary flow. Every record is statistically steady. A drifting flow adds a long-lag correlation that the first zero and the blocks both mistake for a longer integral scale, which is the non-stationarity the essay before warned about and which no estimator here detects.
The convention the numbers depend on
The true integral scale of both processes is one, and every estimate is a fraction of it. Records are sampled at five samples per integral scale except where the sampling rate is the variable. “Error” is the root-mean-square difference from the truth across records, bias and scatter together. The exponential fit uses lags where the sample correlation exceeds 0.3; the block estimate uses the record’s own mean and variance.
Who found it, and when
The first-zero convention is as old as the integral scale, which is G. I. Taylor’s from the 1920s. Estimating an error from the scatter of block means is Bartlett’s method of 1946 for spectra, and its use to measure a correlation time from the plateau of the block variance was popularised in molecular simulation by Flyvbjerg and Petersen in 1989, where records are also long, correlated and expensive. Comparisons of integral-scale estimators on turbulence data have repeatedly found the spread between methods to exceed the scatter of any one, which is the empirical form of the result here.
It sits beside the moment a spectrum cannot hold, where a correlation’s shape at the origin decides whether a spectral moment exists, and the range a real Reynolds number does not have, where a model spectrum’s shape is the assumption doing the work.
Still open: a field instead of a record
Every estimate here reads a time series at one point. A particle-image velocimetry frame or a simulation snapshot gives a field instead, and the integral scale is then a length estimated from spatial correlations across a finite window — the window every vector is averaged over sets the smallest separation and the frame’s size the largest. The calculation that follows repeats the comparison in two dimensions: fields of known integral length generated exactly, the four estimators applied across windows of increasing size, and the question whether a snapshot of given size holds more or fewer independent values than a record of the same number of samples. The answer decides how many snapshots the exact laws’ constants need before their error bars are honest.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- The span takes away the infinity, not the gust — both name averaging, convergence, integral scale, spectrum, statistics
- A dissipation correlated across every scale — both name correlation, measurement, model validity, spectrum
- A drift made of two things that average to zero — both name averaging, correlation, measurement, model validity
- Equal on average, and nothing else — both name averaging, correlation, measurement, model validity
- How far a parcel gets — both name correlation, integral scale, measurement, statistics
- The shutter is part of the answer — both name averaging, measurement, model validity, sampling
Named objects
A dashed tag is an object no other essay names yet.
AveragingConvergenceCorrelationIntegral scaleMeasurementModel validityProbability distributionSamplingSpectrumStatistics