Viscosity

The cheapest shape the walls allow

The parabola in a pipe is usually presented as what the equations give. It is better understood the other way round — of every profile that could carry that flow between those walls, it is the one that destroys the least energy, and the equations give it for that reason.

Worth reading first: The price of a gradient · Mass has nowhere to go.

Every course in fluid mechanics arrives at the parabola the same way. Write the Navier–Stokes equations, throw away everything that vanishes in a long straight channel, integrate twice, apply no slip at both walls. Out comes u=umax(1y2/h2)u = u_{\max}(1 - y^2/h^2), and the reason it is a parabola is that the equation was second order and the source term was constant.

That is a correct derivation and it explains nothing. Here is a different route to the same shape, which explains a great deal.

Of all the velocity profiles that conserve mass, vanish at both walls and carry the required flow rate, the parabola is the one that destroys the least energy. Not one of several; the unique minimum. The equations produce it because it is the cheapest, and the cheapness is the reason rather than a coincidence.

Four profiles carrying the same fluid. Four members of the family, each with no slip at both walls and each carrying exactly the same volume flux. They are drawn on the parabola's own peak velocity, so the areas under them are equal by construction. The flatter ones look easier and are not: flattening the middle steepens the wall, and the cost is the square of the gradient.
Fig. 1 Four profiles carrying exactly the same fluid between exactly the same walls. Each is u = A(1 − |y/h|ⁿ) with A fixed by the flux, so the areas under them are equal by construction and the only thing that differs is the shape. The flatter ones look easier and are not.

The family, and where its minimum is

The quickest way to see it is to give the profile one adjustable parameter and differentiate.

Take u=A(1y/hn)u = A(1 - |y/h|^n). Every member has no slip at y=±hy = \pm h. Fixing the flux fixes AA in terms of nn: A=Q(n+1)/2hnA = Q(n+1)/2hn. The dissipation is μ(du/dy)2dy\int \mu (du/dy)^2 dy, and carrying the algebra through gives

D(n)=μQ22h3(n+1)22n1.D(n) = \frac{\mu Q^2}{2h^3}\cdot\frac{(n+1)^2}{2n-1}.

Differentiate the fraction and the numerator is 2(n+1)(n2)2(n+1)(n-2). There is one turning point in the whole family and it is at n=2n = 2.

The parabola is the cheapest shape the walls allow. The dissipation of the profile u = A(1 − |y/h|ⁿ) carrying a fixed flux between fixed walls, against the exponent n, divided by the parabola's. Every member of the family satisfies the same boundary conditions and carries the same fluid; they differ only in shape. The minimum is at n = 2 exactly, which is not a coincidence — it is Helmholtz's theorem, and the parabola is a solution of the equations because it is the least dissipative shape rather than the other way round.
Fig. 2 The dissipation of the family against its exponent, divided by the parabola’s. The minimum is at two exactly, and the curve on either side of it is not steep — which is the single most important thing about this figure. A shape that is noticeably wrong costs very little more.

Two things about that curve deserve separate attention, and the second is the reason the theorem is not better known.

The first is that n=2n = 2 is not approximate. It is the root of a linear factor, and it does not move if the flux changes, the viscosity changes, the gap changes, or the fluid is replaced with something else. It is a property of the geometry and of nothing else.

The second is that the minimum is flat. A profile with n=1.5n = 1.5 — visibly pointier than a parabola — costs four per cent more. One with n=3n = 3 costs seven. A cosine, which is what a smooth guess produces if asked to draw a plausible channel profile, costs 1.47 per cent more than the true answer, and no measurement anybody would make on a real channel would distinguish them.

What “least” is being taken over

The claim needs its constraints stated precisely, because a variational principle with sloppy constraints is worth nothing.

The comparison is over velocity fields that are divergence-free, that take the same values on the boundary, and — in a channel, where the boundary values are all zero — that carry the same flux. Within that class, the Stokes solution has the least dissipation. Anything outside the class is not being compared: a field that leaks mass is not a competitor, and neither is one that slips at the wall.

This is Helmholtz’s theorem, and its proof is three lines. Write any competitor as the Stokes solution plus a difference w\mathbf{w}, where w\mathbf{w} is divergence-free and vanishes on the boundary. The dissipation of the sum is the dissipation of the solution, plus twice a cross term, plus the dissipation of w\mathbf{w}. The cross term integrates by parts into a boundary integral — which vanishes, because w\mathbf{w} does — plus a volume integral against the Stokes equations, which vanishes because they are satisfied. What is left is

D[uS+w]=D[uS]+D[w]    D[uS],D[\mathbf{u}_S + \mathbf{w}] = D[\mathbf{u}_S] + D[\mathbf{w}] \;\ge\; D[\mathbf{u}_S],

with equality only when w\mathbf{w} is zero. The excess is exactly the dissipation of the difference field.

The consequence that can be measured

That last line is a sharper statement than “the parabola is the minimum”, and it is testable in a way that the minimum itself is not.

If the excess is D[w]D[\mathbf{w}], then scaling the disturbance by ε\varepsilon scales the excess by ε2\varepsilon^2 — with no linear term at all. A functional that had a linear term would have a direction in which the flow could be made cheaper, and there would be no minimum.

The excess is quadratic and has no linear part at all. Add a disturbance of amplitude ε to the parabola, keeping the flux and the walls, and measure what it costs. The excess falls as ε squared over three decades, with a fitted slope of 2.017. That the linear term is missing is the theorem itself: the cross term in the dissipation integrates to zero only because the parabola solves the equations, so a field that did not would show a slope of one and a first-order gain available in one direction.
Fig. 3 The excess dissipation of a disturbed parabola against the size of the disturbance, over three decades, both logarithmic. The fitted slope is two. A field that was not a solution of the equations would show a slope of one somewhere on that range — a first-order gain available in one direction — and there is none.

That is the rejection test this site’s habit asks for. The theorem is not being illustrated; it is being given something it could fail, which is a measured exponent, and the exponent comes out at two.

Where the expensive shapes are

The family above is well behaved because every member of it is smooth. The instructive part of the comparison is what happens when a profile is allowed to be a poor shape rather than a slightly wrong one.

What it costs to be the wrong shape. Every profile here carries the same flux between the same walls, and every one of them costs more than the parabola. A cosine — the shape a smooth guess produces — costs only 1.5 per cent more, which is why the theorem is easy to believe and hard to notice. A plug with a thin shear layer costs without limit, because the price of a gradient is its square and thinning the layer raises the gradient faster than it removes fluid.
Fig. 4 Seven profiles at one flux between one pair of walls, and what each costs above the least. The smooth alternatives are within a few per cent. A plug with all its shear crowded into a thin layer at the wall has no upper bound at all: halving the layer’s thickness roughly doubles the bill.

The plug is the case worth understanding, because it is the shape a turbulent flow actually has.

A plug of speed AA joined to the wall by a ramp of thickness δh\delta h dissipates 2μA2/δh2\mu A^2/\delta h. As the layer thins, the gradient in it rises as 1/δ1/\delta and the volume falls as δ\delta, so the product rises as 1/δ1/\delta — without limit. Concentrating the shear is expensive, and it is expensive by exactly the same arithmetic that makes the wall the only place a pipe’s heat is made.

The same statement in a bearing, where it is worth money

A channel is the clean case and it is not the case anybody is paid to think about. The version that matters commercially is a bearing, and there the theorem says something a designer can use.

A journal bearing carries its load because the film it runs on is converging, and the pressure it generates is a consequence of viscous flow through a narrowing gap. The load capacity is therefore inseparable from the dissipation: a bearing that destroyed no energy would carry nothing. What the theorem adds is that the relation between them is stationary — the film shape that a real bearing settles into is a minimiser, so a small error of form changes the load at second order rather than first.

That is the reason a plain bearing is as forgiving as it is. A pad machined a few per cent off its intended taper does not lose a few per cent of its load; it loses a fraction of a per cent, because the quantity being spoilt is at a minimum with respect to exactly that kind of change. The whole design tradition of running clearances quoted to one significant figure rests on it, and nobody states it.

The best taper is 2.1887, and it is a root rather than a rule of thumb. Load per unit width against the ratio of inlet film to outlet film, for a pad 50 mm long with a 25 µm outlet clearance. At a ratio of one — a parallel film — the load is exactly zero, which is the claim this whole essay is about. It rises to a maximum at 2.1887, found here twice by arithmetic that shares nothing: a golden-section search on the load, and a bisection on its derivative. The two agree to 6e-8, and the curve is flat enough near the top that a bearing built at 2 or at 2.5 loses under two per cent.
Fig. 5 The load a tapered pad carries against its taper ratio. The maximum is broad — a taper anywhere between about 1.7 and 3 gives most of the available load — and the broadness is the same second-order insensitivity the minimum-dissipation theorem guarantees, seen from the other side.

What happens when the constraint is a velocity rather than a flux

One more variation is worth doing, because it changes the answer and is the case a great deal of machinery is actually in.

In the channel above the flux was held and the profile was free. In a Couette flow it is the wall speed that is held, and the flux is whatever it turns out to be. The minimiser is then the straight line, and the argument is the same: any competitor is the straight line plus a field vanishing at both walls, the cross term integrates away, and the excess is the disturbance’s own dissipation.

The straight line, unlike the parabola, dissipates uniformly. Every part of the gap is being sheared at the same rate, so every part is paying the same. That is a striking difference from the pressure-driven case, where nearly all of the bill is at the wall, and the two cases sit side by side in almost every real film — a bearing is driven by both a wall speed and a pressure gradient at once, and the dissipation it makes is not the sum of the two considered separately.

Where a pipe's heat is made. Plane Poiseuille flow: the velocity profile on the left, the dissipation function on the right, at the same scale of height. The fluid in the middle is moving fastest and is dissipating nothing at all, because it is not being sheared; every joule is made at the walls, where the fluid is barely moving. Half of the total is made in the outer 29 per cent of the gap.
Fig. 6 The pressure-driven case, with its dissipation piled against the walls. The wall-driven case, drawn on the same axes, would be a flat line — and comparing the two is the quickest way to see that “where a film’s heat is made” is not a property of the film but of what is driving it.

Which is why the theorem does not say what it seems to

Here is the trap, and it is worth setting out plainly because the theorem is frequently quoted with the constraints dropped.

Turbulent pipe flow has a plug-like profile with a thin wall layer, which is the shape the law of the wall describes. By the arithmetic above it must dissipate far more than the parabola at the same flow rate — and it does, by roughly a factor of three at Reynolds number 10510^5, rising with Reynolds number because the two friction laws have different exponents.

So a real flow at that Reynolds number is emphatically not minimising its dissipation. It is doing several times worse than a solution that exists, satisfies the same boundary conditions, and carries the same flux.

There is no contradiction, and locating it precisely is the useful exercise. Helmholtz’s theorem is a statement about the Stokes equations, in which the nonlinear term is absent. Those equations have exactly one solution for given boundary data, so “the solution” and “the minimiser” are the same object and there is nothing for a flow to choose between. The Navier–Stokes equations have no such guarantee: the laminar solution is still a solution, and it is no longer the only one, and nothing in the equations says the cheapest available one is taken.

Minimum dissipation is a property of uniqueness, not a principle of selection. Every attempt to turn it into one — and there have been many, from Helmholtz onwards — founders on the same fact.

The variational idea, running the other way

Minimum dissipation is not a selection principle is the right conclusion and it is not the end of the variational programme, because the same style of argument survives if the inequality is turned round.

The failure above is that the laminar solution is a lower bound the flow declines to take. Ask instead for an upper bound — how much can a flow between these walls dissipate, whatever it is doing — and the question becomes answerable, and the answer is a theorem about the full Navier–Stokes equations with no closure, no model and no measurement in it.

The modern machinery is disarmingly simple in outline. Split the velocity into a steady background field that carries the boundary conditions, plus a fluctuation that vanishes on the walls. Substituting that split into the energy balance leaves a quadratic form in the fluctuation, and if the background is chosen so that the form is positive definite, everything the fluctuation could possibly be doing is bounded — so the dissipation is bounded, by a quantity depending only on the background. Optimising over backgrounds gives the best bound the method can produce.

What comes out is genuinely a theorem. For shear-driven turbulence it gives a drag coefficient bounded independently of Reynolds number, which is the rigorous half of the dissipation anomaly: the statement that the losses do not fall as the viscosity does, obtained without assuming anything about eddies. For a heated layer it gives a Nusselt number bounded by the square root of the Rayleigh number.

The bounds are not tight — typically a factor of several above what is measured — and that is the honest position. What the variational method cannot do is say what a turbulent flow does; what it can do is prove what no flow may exceed, which is the same relationship this collection’s control-volume field has to the machines it bounds.

What it is good for anyway

Three things, and all of them are about bounding rather than predicting.

It bounds a drag from above. Any admissible field’s dissipation is an upper bound on the true one, so guessing a plausible profile and integrating gives a number that is too large and is known to be too large. That is how the first estimates of the drag on awkward shapes were made, and it is why they were quoted as bounds rather than as answers — including the estimates that preceded the exact solutions this collection uses for a sphere.

It says perturbations are quadratic. A slightly deformed geometry has a dissipation that differs from the original at second order in the deformation, not first — which is why a bearing’s load capacity is insensitive to small errors of form and why the harmonic mean sets the pressure peak of a tapered pad so robustly.

It explains why a Stokes solve is forgiving. A numerical Stokes solve that is slightly wrong is wrong in its dissipation by the square of how wrong it is in its velocity. That is a genuine practical advantage and it has no counterpart at high Reynolds number: a Navier–Stokes solve that is one per cent wrong in its velocity field can be several per cent wrong in its drag, because there is no stationarity to protect it.

45° of corner, and an infinite number of eddies in it. The creeping flow in a corner of 45 degrees, drawn from Moffatt's similarity solution. Each eddy turns the opposite way to its neighbours and is 3.17 times smaller and 1.6e+3 times weaker than the one outside it. The contour levels are rescaled inside each eddy, because they differ in strength by three orders of magnitude per step and a single set of levels would show the first and nothing else — which is itself the reason nobody has seen the third. The dividing lines between eddies are drawn where the stream function changes sign, and the dots are the centres, both found from the solved field.
Fig. 7 An unexpected consequence: the infinite sequence of ever-smaller eddies in a sharp corner. Each is a Stokes flow driven by the one above it, each is the least-dissipating field consistent with the boundary the one above provides, and the whole cascade is a minimum-dissipation solution all the way down — the same eddies this collection computes an eigenvalue for. Nothing is choosing to make eddies; they are what costs least.

What the picture cannot show

The comparison is at fixed boundary values, and a real design changes them. Every profile here carries the same flux between the same walls. A designer who is allowed to change the gap, or to add a wall, is not in the class the theorem covers, and the theorem has nothing to say about that comparison.

The family is one-parameter and the theorem is not. Finding the minimum of D(n)D(n) shows that the parabola beats every member of one particular family. The theorem says it beats every admissible field, which is an infinite-dimensional statement, and the perturbation ladder is the closest this essay comes to testing it in that generality.

And nothing here is about stability. A minimum of the dissipation is not a stable state; the laminar profile is both the global minimiser and, above transition, unstable. Those are answers to different questions and the second is where the neutral curve is.

The same number, by two integrals that share no arithmetic. Three flows whose dissipation is in closed form both ways. The volume route integrates the dissipation function over the fluid; the boundary route multiplies a force or a torque by the speed of whatever is applying it. Neither calculation contains the other, and the residual column is what is left when they are subtracted.
Fig. 8 The calibration this whole family of results rests on. If the dissipation of a given field is being compared with another field’s, both integrals had better be right, and the check for that is the one this collection uses everywhere: compute the same number two ways that share no arithmetic.

The number the flatness explains

The flatness of the minimum has one more consequence, and it is the reason this theorem is more often useful as an excuse than as a tool.

Almost every practical estimate of a viscous flow is made by assuming a profile. Lubrication theory assumes a parabola across the film; integral boundary-layer methods assume a one-parameter family; network models of a piping system assume fully developed flow in every branch. Each of those assumptions is wrong in detail, and each of them is wrong in a direction that raises the dissipation, because the true profile is the minimiser.

So the errors are all of one sign and all of second order. An assumed profile that is ten per cent wrong in shape gives a dissipation that is about one per cent too high — never too low, and never proportionally. That is why lubrication theory works as well as it does on films that are not quite thin, and why an integral method’s drag is usually good to a few per cent while its profile is visibly not the right shape.

The exception is the shape that is wrong in the expensive direction. A method that assumes a plug where the flow is really parabolic is not making a second-order error; it is on the steep part of the curve, and the further it goes the worse it gets without bound. That asymmetry is worth carrying: the penalty for guessing a smooth profile is negligible, and the penalty for guessing a flat one is not.

Who found it, and when

Helmholtz proved it in 1868, and Korteweg gave the converse — that the minimiser satisfies the Stokes equations — in 1883, which is the half that makes it a genuine variational principle rather than a property. Rayleigh restated it in 1913 in the form most often quoted, and spent some effort trying to extend it to flows with inertia, which cannot be done.

The surprising connection is with electrical networks. The dissipation of a Stokes flow is a positive-definite quadratic form in the velocity field, minimised subject to a linear constraint — which is the same mathematical object as the power dissipated in a resistor network, minimised subject to Kirchhoff’s current law. Thomson’s principle for currents and Helmholtz’s for Stokes flows are the same theorem about the same kind of functional, discovered independently fifteen years apart, and both fail for exactly the same reason when the medium stops being linear.

Where the ladder goes next

Beside this rung is the reciprocal theorem, which is the other thing linearity buys: a force obtained without solving for the flow that makes it. Both are properties of the Stokes equations and both stop at the first appearance of inertia.

Below it is the price of a gradient, which supplies the functional being minimised, and mass has nowhere to go, which supplies the constraint. Above it, in a sense, is what it costs to go turbulent — the measurement of how far a real flow sits from the cheapest one available to it.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Boundary conditionDissipationIrreversibilityMinimum-dissipationOptimisationPoiseuille flowStokes flowTurbulenceVariational principleViscosity