Regimes and numbers

The groups are not the only groups

Buckingham's theorem fixes how many dimensionless groups an answer can depend on and says nothing about which. Two of the infinitely many legitimate choices are used here on the same data: one manufactures a straight line through five decades out of a constant, and the other erases Stokes' law completely.

Worth reading first: Counting what matters · Exact in the total, free in the profile.

Buckingham’s theorem is a rank. Five quantities decide the drag on a sphere, their dimension matrix has three independent rows, and what is left over — two — is how many dimensionless groups the answer can possibly depend on.

That much this collection has already worked out rather than quoted. What it stopped at, as almost every account of the theorem stops, is the count.

The dimension matrix, and the rank that is the whole theorem. Five quantities decide the drag on a sphere and their dimensions fill a three-by-five matrix. Its rank is three, so its null space has dimension two, and that is Buckingham's theorem: the answer can depend on two dimensionless groups and no more. The rank is unique. Nothing in the theorem says which two.
Fig. 1 The dimension matrix, and the rank that is the whole theorem.

The rank is unique and the basis is not. The null space of that matrix is a two-dimensional lattice, and any pair of integer vectors generating it is as valid a set of pi groups as any other. There are infinitely many, related by the integer matrices of determinant plus or minus one, and the theorem has no preference among them. The drag coefficient and the Reynolds number are a convention.

The lattice is worth a moment, because it is what makes “infinitely many” precise rather than rhetorical. A dimensionless combination of the five variables is an integer vector of exponents lying in the null space of the dimension matrix, and those vectors form a two-dimensional grid of points — a lattice. A set of groups is a pair of grid points from which every other can be reached by integer steps. There are infinitely many such pairs, exactly as there are infinitely many pairs of vectors generating the integer points of a plane, and the theorem’s content is the dimension of the grid rather than any particular pair.

The integer part matters and is easy to lose. Computing the null space over the rationals and clearing denominators at the end — the obvious thing — returned, for this problem, the pair ρU2d2/F\rho U^2 d^2/F and μ2/Fρ\mu^2/F\rho. Both are perfectly dimensionless. Both are in the null space. And they generate only half of it: the Reynolds number is not an integer combination of them, and the pair they form has determinant two against the true basis. The failure is silent, because every test one would apply to a pi group passes on both.

Four sets of groups, all of them complete

Four sets of groups, all of them complete. The computed null-space basis and three recombinations of it by integer matrices of determinant one. Every pair is dimensionless, every pair generates the whole lattice, and the theorem prefers none of them. The last row is a recombination of determinant two, whose groups are perfectly dimensionless and generate only half the lattice — a defect nothing about the groups themselves reveals.
Fig. 2 Four sets of groups, all of them complete.

Take the computed basis and act on it with an integer matrix of determinant one. The result is another pair of dimensionless numbers, and the pair generates the same lattice — which is what makes it a complete set, meaning that every dimensionless combination of the five variables is a product of powers of those two. All the recombinations in the table are complete sets. All of them are legitimate. A paper written in any of them is a correct paper.

The last row of that table is the one worth pausing on. It is a recombination of determinant two. Its two groups are perfectly dimensionless — every check anybody applies to a pi group passes on them — and they generate only half the lattice. What has gone is completeness, and nothing about the groups themselves reveals that: a dependence can hide in the quotient, and the count of groups is still two.

What a recombination can do to data

The freedom would be an idle observation if every basis told the same story. Two of them do not.

A drag coefficient that does not depend on the Reynolds number. Two hundred and forty points of synthetic data in the Newtonian regime, where the drag coefficient is 0.44 with twelve per cent scatter and no Reynolds-number dependence whatever. The physical variables behind each point are drawn at random over decades, so a collapse would be a collapse. The fitted slope is −0.0015 at an r-squared of 0.0015.
Fig. 3 A drag coefficient that does not depend on the Reynolds number.

Here is a data set in the Newtonian regime: two hundred and forty measurements in which the drag coefficient is 0.44 with twelve per cent scatter and no Reynolds-number dependence at all. The density, viscosity, diameter and speed behind each point are drawn independently over decades, so a collapse would be a collapse rather than a sweep of one variable. The fitted slope is −0.0015 at a coefficient of determination of 0.0015: there is nothing there, and the plot says so.

The same data, in a basis the theorem allows just as much. The identical two hundred and forty points, plotted as F/(mu U d) — which is the drag coefficient times the Reynolds number, a perfectly legitimate pi group forming a complete pair with the abscissa. The result is a straight line of slope 0.9985 through five decades with an r-squared of 0.9985, and it contains no physics: the ordinate contains the abscissa.
Fig. 4 The same data, in a basis the theorem allows just as much.

Now the identical points, plotted as F/μUdF/\mu U d against the Reynolds number. That group is CdReC_d \mathrm{Re}; it is dimensionless; it forms a complete pair with the abscissa; the recombination matrix has determinant one. And the result is a straight line of slope 0.9985 through five decades with a coefficient of determination of 0.9985.

There is no physics in it whatever. The ordinate contains the abscissa, so the plot is the abscissa against itself with a constant of proportionality and a little scatter. A reader shown it would conclude that the force rises in proportion to the Reynolds number, which in this data set it does not do at all.

It is worth being clear how ordinary the offending group is. F/μUdF/\mu U d is a force divided by a viscous force scale, which is a perfectly natural thing to form and is exactly the group somebody working in the world with no inertia would reach for first — there the viscous scale is the only one available and the drag coefficient is the artificial choice. Nothing about the group is a trick. What makes the plot empty is the pairing: it is on the same axes as a group that carries the same viscosity, the same speed and the same diameter.

And the same recombination erases Stokes' law. Data in the creeping-flow regime, where the drag coefficient is 24/Re — as strong a dependence as this subject has. Plotted conventionally the slope is −1.0015 at an r-squared of 0.9986. Plotted in the recombined basis it is flat, at a slope of −0.0015 and an r-squared of 0.0015. A reader shown the second would conclude the force does not depend on the Reynolds number at all.
Fig. 5 And the same recombination erases Stokes’ law.

The reverse is worse. Take creeping flow, where Cd=24/ReC_d = 24/\mathrm{Re} — as strong a dependence as this subject has, and the one every account of low Reynolds number opens with. Plotted conventionally the slope is −1.0015 at an r-squared of 0.9986. Plotted in the recombined basis it is flat: a slope of −0.0015 at an r-squared of 0.0015.

A reader shown the second figure would conclude that the force does not depend on the Reynolds number. The same recombination that manufactured a law out of nothing has erased one that was there, and it did both while satisfying every requirement Buckingham’s theorem imposes.

The erasure has a use, which is the part that keeps this from being a cautionary tale about a bad choice. CdReC_d \mathrm{Re} being constant is Stokes’ law, stated as compactly as it can be: one number, 24, with no residual dependence in it. Anybody working in that regime is right to use the group, and the flat line is the result rather than the absence of one. What makes the same figure a failure in the Newtonian regime and a success in the creeping-flow one is not the group. It is whether the flatness is the physics or the arithmetic, and the plot cannot tell a reader which.

Four fits, two data sets, one theorem. The same two collections of points, each seen in the conventional basis and in a legitimate recombination of it. A constant becomes a perfect straight line and a strong power law becomes a horizontal one, and every group in every row is dimensionless and part of a complete set.
Fig. 6 Four fits, two data sets, one theorem.

What is actually invariant

None of this means the collapse itself is a convention. It is not, and separating the two is the point of the next figure.

Two physical states are built with the same Reynolds number and nothing else in common: a two millimetre bead in air, and a forty centimetre boulder in glycerine whose density is a thousand times higher and whose viscosity is eighty thousand times higher. Every group in every basis agrees between them, to the last bit of double precision, at six Reynolds numbers from 0.1 to 5,000.

That is what a complete set of groups means, and it is the theorem’s real content: the physical variables have nothing in common and the groups have everything.

Sphere drag, seen in four complete sets of groups. One data set, four bases. Every one of them collapses, because every one of them is a complete set and the physics has only two groups in it. What changes is what the curve looks like — an exponent, a slope, a curvature — and therefore what a reader takes away from it.
Fig. 7 Sphere drag, seen in four complete sets of groups.

So the data collapse in every basis — they must, because the physics has two groups in it and every complete set is a coordinate system on the same two-dimensional space. What changes between bases is what the curve looks like: an exponent, a slope, a curvature, a flatness. And what a reader takes away from a plot is its shape.

This separates two things that are often run together. The collapse is the theorem’s promise and it is kept: a bead and a boulder at one Reynolds number are one point, and no basis breaks that. The shape of the curve is not the theorem’s promise and never was, and it is the shape that carries the exponents people quote, the asymptotes they extrapolate along and the transitions they name. A choice with no consequence for the collapse can have every consequence for those.

It also puts a familiar practice in an awkward light. Sweeping a single variable — running one sphere at several speeds — produces points that collapse in any basis, complete or not, because every group varies together along the sweep. A collapse from such an experiment is not evidence of anything at all, which is why the data here are drawn independently across every variable. It is also why a model that cannot be matched is a real difficulty rather than a pedantic one: matching every group at once is what makes a collapse mean anything, and a model test that matches one group and sweeps another has produced a curve whose shape is partly the sweep.

The rule that is not in the theorem

There is a rule that separates the honest plots from the artefacts, and it is worth stating plainly because it cannot be derived from anything above.

The quantity being predicted must appear in exactly one group, and not in the one it is plotted against.

The rule the theorem does not contain. Buckingham's theorem fixes the number of groups and nothing else. What decides whether a plot means anything is a rule from outside it: the quantity being predicted must appear in exactly one group, and not in the one it is plotted against. Every failure on this page is a violation of that rule by a set of groups the theorem itself has no objection to.
Fig. 8 The rule the theorem does not contain.

The drag coefficient satisfies it: the force appears once, in the ordinate, and the abscissa is built from the four variables that were set rather than measured. CdReC_d \mathrm{Re} does not: the force is still only in the ordinate, but the ordinate now also carries the viscosity, the speed and the diameter that the abscissa carries, and the shared factors are what produce the line.

That is a version of a hazard which is not peculiar to fluid mechanics. Plotting A/BA/B against C/BC/B with AA, BB and CC independent gives a strong correlation from the shared BB alone, and the statistics literature has called it spurious correlation since Karl Pearson named it in 1897. What dimensional analysis adds is that the shared factors arrive disguised as physics, wearing the names of legitimate dimensionless groups, with a theorem apparently endorsing them. The pi groups are not the culprit — a pure number does come from somewhere, and where it comes from is a solution rather than a plot.

The other choice nobody records

The textbook recipe does not usually construct a null space. It picks three repeating variables, forms each group from those three plus one other, and requires the three to be dimensionally independent.

One repeating set out of ten does not work. The textbook recipe forms the groups from three repeating variables plus one other, and the three have to be dimensionally independent. Nine of the ten triples available here are; the tenth — force, density and viscosity — has a singular sub-matrix, and its determinant is exactly zero rather than small.
Fig. 9 One repeating set out of ten does not work.

Nine of the ten available triples here are usable. The tenth — force, density and viscosity — has a singular sub-matrix: its determinant is exactly zero, not small, and no set of groups can be built from it. That is a failure the recipe catches.

What the recipe does not catch is the choice among the nine that work. Repeating on density, speed and diameter gives the drag coefficient and the Reynolds number. Repeating on density, viscosity and diameter gives the pair that produced the straight line above. Both are correct applications of the same instructions, and only one of them produces a plot with information in it.

The reason the recipe cannot help is that it is a procedure for finding a basis, and the trouble is in choosing one. Its single test — dimensional independence of the repeating variables — is a test of whether the construction will succeed, not of whether the result is worth plotting. Which is why the failure survives every check a careful reader would apply: the groups are dimensionless, the count is right, the set is complete, and the plot is empty.

A second problem, to show the counting is not about spheres

Six quantities and three dimensions leave three groups, which is why a Moody chart is a family of curves and a sphere-drag chart is one. The rank does the work; the names on the axes do not.

And the same freedom is there. A friction factor plotted against a Reynolds number formed from the same pressure gradient the friction factor contains would collapse beautifully, for the same reason, and would be exactly as empty.

Three groups also make the choice harder rather than easier, because the freedom is now a three-by-three lattice and the number of legitimate bases is much larger. It is the reason a relative roughness is plotted as a parameter on a family of curves rather than folded into either axis: keeping it separate is a decision about presentation that the theorem does not make and that decides what the chart shows.

There is a related trap in the other direction, and this collection has met it. An exponent that dimensions cannot give is one that has to come from the solution rather than from the counting — and it will look, on a log-log plot in any basis, exactly like an exponent that came from the counting. The theorem constrains the number of groups. It never constrained the exponents at all.

Where the choice is already made for everybody

It is easy to read all of this as a warning about carelessness. It is not: the conventional groups of this subject were chosen well, and the reason they were chosen well is worth recovering.

The Reynolds number puts the four variables that an experimenter sets into the abscissa and the one that is measured into the ordinate. That is not a dimensional statement, it is a statement about the experiment, and it is what makes a drag chart readable — every point on it is one measurement, and moving along the axis means changing the apparatus rather than the result.

Which speed goes into the number is the same decision one level down, and this collection has already found that it is not free either: the group is fixed as a form and the velocity scale put into it is a choice with consequences. So is the length. The freedom this essay is about does not stop at the choice of basis; it runs all the way down to which velocity and which length are used inside a group whose name everybody agrees on.

The Mach number, the Froude number and the Weber number are all built the same way, and all of them put the imposed quantities in the group and leave the answer outside it. That consistency is why the conventions of this subject compose — a chart in one can be read beside a chart in another — and it is a convention rather than a consequence.

What to do about it

Three habits, and none of them requires anything beyond the arithmetic above.

Write the null space out, rather than picking repeating variables. It takes one integer elimination, it produces the whole lattice rather than one member of it, and it makes the choice visible as a choice.

Check where the unknown appears. If the quantity being predicted is in both axes, the plot will correlate whether or not the physics does, and the correlation coefficient is measuring the arithmetic.

Vary the physical variables independently, not along a sweep. A collapse produced by changing one thing is a collapse of a one-parameter family and would happen in a wrong basis too.

And read a strong collapse as a question rather than as an answer. An r-squared of 0.9985 across five decades is either a real law or a shared variable, and telling them apart takes ten seconds with the exponents in front of one.

This is the same discipline the rest of this collection applies to a figure: a result that follows from a general theorem is exactly as strong as the theorem, and no stronger. Buckingham’s theorem is a statement about the rank of a matrix. It is true, it is useful, and it endorses nothing about which two of infinitely many groups a figure should be drawn in.

What a correlation is, when it is honest

It is worth ending on what a good dimensionless plot is doing, since so much of this essay has been about what a bad one is doing.

A correlation of drag coefficient against Reynolds number is a measurement of a function. The theorem says the function exists and has one argument; the experiment measures it; and the resulting curve is a compressed statement of every experiment that could have been done, because any state with that Reynolds number has that drag coefficient.

That is an enormous compression and it is the whole value of the method. Five variables that could have needed a five-dimensional table need one curve, and the curve transfers between fluids, sizes and speeds that share nothing.

What the arithmetic on this page adds is a condition on the compression being readable. The curve has to be plotted in coordinates where the thing being predicted appears once, because otherwise the curve is partly a picture of its own axes. That condition is not in the theorem, it costs nothing to check, and it is the difference between a compressed measurement and a compressed tautology.

And there is a corollary about reporting. The exponents in a fitted power law belong to the basis they were fitted in, so quoting one without the groups is quoting half of a statement. A drag falling as the minus one power of the Reynolds number and a force rising as the zeroth power of it are the same measurement, and neither number means anything without the axes.

What is not claimed

The theorem is not being questioned. The rank is three, the number of groups is two, and every recombination used here is exact integer arithmetic on the null space. Nothing on this page is a numerical result about dimensional analysis; the arithmetic is exact and the failures are failures of interpretation.

The synthetic data are synthetic. They are generated from a stated law with stated scatter so that the true dependence is known, which is the only way to demonstrate a spurious one. Real sphere drag does depend on Reynolds number, and the correlation used for the collapse figure is a fit to real measurements rather than anything derived here.

The placement rule is a convention with a reason, not a theorem. There are cases where the predicted quantity legitimately appears in more than one group — a lift coefficient and a moment coefficient share the dynamic pressure, and nobody is misled — and the rule is a guard against a particular failure rather than a law.

And the count of usable repeating sets is a property of this problem. Five variables and three dimensions give ten triples; a different problem gives a different count, and the fraction that fail is not a general number.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Buckingham's pi theoremCorrelationDimensional analysisDimensionless numberDrag coefficientDynamic similarityMeasurementMisconceptionNull spaceRankReynolds numberScaling