Sunday, 22 July 2018

visible light - What determines the apparent radius of the rainbow?


Let's say I know how to compute the apparent radius of a rainbow from the viewpoint of the observer: take a photo of the scene, measure the distance to a known reference object, and its dimensions. Using triangle similarity, I can extrapolate the radius of the rainbow.


But my question is: which physical phenomenon determines the radius?



Answer



It depends on where the sun is. If it is near the horizon (behind you) and in front of you there are water droplets, then you will see a rainbow with a radius (in angular measure) of about 42 degrees, because each water droplet returns a cone of light, whose axis is parallel to the direction to the sun and whose aperture is roughly $2 \cdot 42 = 84$ degrees.


I've never seen better explanations of dozens of phenomena concerning rainbows than in Walter Lewin's lectures.


Saturday, 21 July 2018

quantum field theory - Gauge fermions versus gauge bosons


Why are all the interactions particle of a gauge theory bosons. Are fermionic gauge particle fields somehow forbidden by the theory ?



Answer



The reason that the gauge particle must be a spin 1 gauge boson is because there aren't any renormalizable alternatives. To see this consider the Dirac Lagrangian:


\begin{equation} \bar{\psi} i \gamma ^\mu \partial _\mu \psi \end{equation} This term is not gauge invariant under the transformation, $ \psi \rightarrow e ^{ i T ^a \theta ^a (x) } \psi $, because of the derivative spoils the desired transformation of $ \partial _\mu \psi $. To fix this we must add a contribution that transforms in the same way as the derivative, i.e., transforms as a vector. In other words we modify the derivative such that, \begin{equation} D _\mu \psi \rightarrow e ^{ i T _a \theta _a (x) } D _\mu \psi \end{equation}



The question is what to add to $ D _\mu $. We can potentially add spin $ 0, \frac{1}{2} , 1 , \frac{3}{2} , $ and $ 2 $ particles to fix this. We go case by case.


There is no combination of spin zero fields that transform as a vector without adding derivatives (adding derivative to fix the derivative covariance would take you in circles), thus we can't have a spin zero gauge boson.


Next consider adding a spin $ 1/2 $ gauge boson we could write ($ \psi _a $ is a gauge particle, not $ \psi $), \begin{equation} D _\mu = \partial _\mu + \sum _a T ^a \left( g\bar{\psi} ^a \gamma _\mu \psi ^a + g ' \bar{\psi} ^a \gamma _\mu \gamma ^5 \psi ^a \right) \end{equation} However, this would give an interaction \begin{equation} \sum _a i\left[ \bar{\psi} \gamma ^\mu\psi \right] \left[ \bar{\psi} _a \gamma _\mu \psi _a \right] \end{equation} and similarly for the $ \gamma ^5 $ term. These interactions are non-renormalizable as they involve four fermions. Non-renormalizable interactions arise from effective field theories and are suppressed by the scale at which they arise. This would make the gauge interactions non-fundamental but instead involve a massive vector particle integrated out. For the integrated out interaction to be renormalizable it must be between two fermions and a spin $1$ field. This brings us back to the usual case.


The spin $1$ field works well and exists in the SM. I'm not sure about the spin $ 3/2 $ field as I have no experience with working with such fields however, I presume it won't work for similar reasons. I also know that spin $2$ fields must mediate gravitational fields and thus would give a nonsensible result.


special relativity - Can I fix a point in Minkowski space to give it a vector space structure?


I looked up the term Minkowski space on Wikipedia. It said



There is an alternative definition of Minkowski space as an affine space which views Minkowski space as a homogenous space of the Poincaré group with the Lorentz group as the stabilizer.



In their book Metric Affine Geometry, Snapper and Troyer state on page 59:



It cannot be stressed enough that the affine space $X$ is not a vector space. Its points cannot be added and there is no way to multiply by scalars. No point in $X$ is preferred; they all play the same role. In particular, there is no point in $X$ which makes a better origin for a vector space than any other point.


The situation changes radically if we choose a point $c$ in $X$ and keep it fixed. It is now possible to make $X$ into a left vector space over $k$ by using the one-to-one mapping $f$ from $X$ onto $V$ defined by $f(x) = \overrightarrow{c,x}$ for each $x \in X$. All we do is carry the vector space structure of $V$ over to $X$ by means of the mapping $f$.


So here's my question: As I understand it, it makes sense to think of Minkowski space as an affine space since the basic principle of Special Relativity is that no point in $X$ is a preferred reference frame. But does that mean it is then impossible to "fix" a point in $X$ as Snapper and Troyer say can be done? In other words, is there any physical meaning to the idea of fixing a point in the affine space or is that impossible according to SR?


Obviously I am trying to use a mathematician's idea to interpret what can be done physically with Minkowski space.



Answer



Fixing a point is more or less like fixing a coordinate system on your affine space. Then you can identify $X$ with $V$ as stated in the book, where the fixed point $c\in X$ is mapped to the origin of $V$. In other words, fixing a point $c$ in $X$ is like glueing a copy of $V$ onto $X$ in such a way that $O\in V$ overlaps with $c\in X$. As far as the Lorentz group is considered then the coordinates (i.e. the component of the glued copy of $V$ onto $X$) really behave like vectors, but this is no longer the case under more general transformations (consider for instance translations, or the action of the ray inversion from the conformal group).


radiation - What does the decay constant mean?


In my curriculum, the decay constant is "the probability of decay per unit time"


To me, this seems non-sensical, as the decay constant can be greater than one, which would imply that a particle has a probability of decaying in a time span that is greater than 1.



Can someone explain this?



Answer



You're missing two things. First, that the decay constant is the probability of decay per unit time. That part is important. The actual decay probability over a short time period is equal to the probability per unit time, multiplied by the time period:


$$P = \lambda\Delta t$$


$\lambda$ can be as large as you like, but for a small enough interval $\Delta t$, you'll still have $P < 1$. So there's no contradiction there.


The other thing you're missing is that $\lambda$ is only the probability per unit time given that the nucleus has not already decayed. That's also important. You have to start with an undecayed nucleus.


So let's say you have an undecayed nucleus at $t = 0$.


\begin{align} P_0(\text{decayed}) &= 0 & P_0(\text{undecayed}) &= 1 \end{align}


After some short time $\Delta t$, the probability that it will have decayed is $\lambda\Delta t$, as above.


\begin{align} P_1(\text{decayed}) &= \lambda\Delta t & P_1(\text{undecayed}) &= 1 - \lambda\Delta t \end{align}



Now consider the next time interval, from $t = \Delta t$ to $t = 2\Delta t$. If the nucleus didn't decay in the first time interval, it has a probability $\lambda\Delta t$ of decaying in this second interval. But if the nucleus did decay in the first time interval, the probability that it will have decayed by the end of the second time interval is 1. So overall, the probability that it has decayed by $t = 2\Delta t$ is


\begin{align} P_2(\text{decayed}) &= P_1(\text{undecayed})\lambda\Delta t + P_1(\text{decayed})(1) \\ &= (1 - \lambda\Delta t)\lambda\Delta t + \lambda\Delta t \\ &= (2 - \lambda\Delta t)\lambda\Delta t \\ P_2(\text{undecayed}) &= P_1(\text{undecayed})(1 - \lambda\Delta t) \\ &= (1 - \lambda\Delta t)^2 \end{align}


You can probably see the pattern from here:


\begin{align} P_3(\text{decayed}) &= P_2(\text{undecayed})\lambda\Delta t + P_2(\text{decayed})(1) \\ &= (1 - \lambda\Delta t)^2\lambda\Delta t + (2 - \lambda\Delta t)\lambda\Delta t \\ &= \bigl(3 - 3\lambda\Delta t + (\lambda\Delta t)^2\bigr)\lambda\Delta t \\ P_3(\text{undecayed}) &= P_2(\text{undecayed})(1 - \lambda\Delta t) \\ &= (1 - \lambda\Delta t)^3 \end{align}


In particular, at $t = n\Delta t$,


$$P_n(\text{undecayed}) = (1 - \lambda\Delta t)^n$$


Now, in the limit where $\Delta t$ is short, and $n$ is large, as it must be if $T = n\Delta t$ is going to be a normal-scale time interval, you may recognize this as an exponential:


$$\lim_{n\to\infty}P_n(\text{undecayed}) = \lim_{n\to\infty}(1 - \lambda\Delta t)^n = \lim_{n\to\infty}\biggl(1 - \frac{\lambda T}{n}\biggr)^n = e^{-\lambda T}$$


So the equation for exponential decay emerges naturally from the fact that the decay constant is the decay probability per unit time for an undecayed nucleus. (Or of course the same argument applies to any other system that undergoes exponential decay, not just nuclei.)


homework and exercises - A big cannon to match a ballistic missile?



During the last years of WW2 the Germans used ballistic missiles V-2 (with payload mass ~1,000 kg) to bombard London, from a distance about 300 km away. Suppose the British could respond by building a cannon of a very large size, to shoot back conventional projectiles. Would this be realistic? Neglecting the air, the maximum projectile range (shot at 45 degrees to the horizon) is $R=V^2/g$, where $V$ is the projectile muzzle velocity. For $R$=300,000 m and $g$=10 m$^2$/s one needs $V$=1,700 m/s. This is on the high side but within the reach for conventional rifles and cannons. Does the presence of the Earth atmosphere make it practically impossible to shoot projectiles at such large distances? Since the air resistance should scale with the projectile cross-section, $F \sim \rho V^2 A$, can a very long and thin hypersonic projectile travel easily through the atmosphere?





How many physical degrees of freedom does the $mathrm{SU(N)}$ Yang-Mills theory have?


The $\mathrm{U(1)}$ QED case has two physical degrees of freedom, which is easy to understand because the free electromagnetic field must be transverse to the direction of propagation. But what are the physical degrees of freedom the $\mathrm{SU(N)}$ Yang-Mills theory?


I believe that the $\mathrm{SU(N)}$ Y.M. theory has four degrees of freedom, and it is easy to see that by gauge fixing (such as the Lorentz condition or the Coulomb condition) that we can always remove one redundant degree of freedom. But then we still haven't completely fixed the redundant degrees of freedom. Thus my question is:



How many physical degrees and how many are redundant degrees of freedom does the Y.M. theory have?



I would be interested to understand how we can determine this mathematically, and also to understand what the physical intuition behind this is?



Answer



Since I am new here and I cannot comment above I just write my comment as a small answer. I think in your question there is a small mistake. You ask how many degrees of freedom are there in a Yang-Mills theory although you want to ask how many polarizations does the gluon have. First of all, the number of generators of the gauge group will be the number of gauge bosons you will get. $SU(N)$ groups have indeed $N^2-1$ generators, except $U(1)$ that has one. In QED this corresponds to the photon, in $SU(2)$ to three gauge bosons (not exactly the W's and the Z since these are produced by mixing $SU(2)_L$ with $U(1)_y$), and in $SU(3)$ there are 8 gluons as Siva already told you.



Now to your question, how many polarizations do these gauge bosons have. Let us start with a massive vector boson. We can define its helicity states at the rest frame and then of course boost to the frame that the particle moves. Under this boost, the longitudinal polarization has a term $E/m$ (where $E$ the energy and $m$ the mass of the particle) that goes to infinity when the mass of the particle goes to zero. This would cause problems concerning the unitarity of the theory since the contrubutions in the matrix elements would be huge. What one can do is arrange the interactions of the theory such that the contributions from longitudinal polarizations are suppressed by a factor of the order $m/E$. In the case now the vector boson is strictly massless, the longitudinal contributions decouple completely. Therefore, massless vector bosons have two physical states of maximal helicity while massive one have three. Gluons are massless and so are photons, while W's and Z are massive.


I hope I helped. I am pretty sure that you can find more details even in wikipedia if you want.


differentiation - Total and partial derivatives in thermodynamics and Maxwell relations



Consider the expression $$dS=\left(\frac{\partial S}{\partial T}\right)_VdT+\left(\frac{\partial S}{\partial V}\right)_TdV$$


I'm trying to understand how to derive an expression for $\left( \frac{\partial S}{\partial V} \right)_P$ and how is it related to $\left( \frac{\partial S}{\partial V} \right)_T$.


I tried the following:


Method 1


i) Divide both sides by dV $$\frac{dS}{dV}=\left(\frac{\partial S}{\partial T}\right)_V\frac{dT}{dV}+\left(\frac{\partial S}{\partial V}\right)_T\frac{dV}{dV}$$


ii) and at const. P


$$\left(\frac{dS}{dV}\right)_P=\left(\frac{\partial S}{\partial T}\right)_V\left(\frac{dT}{dV}\right)_P+\left(\frac{\partial S}{\partial V}\right)_T\left(\frac{dV}{dV}\right)_P$$


$$\left(\frac{dS}{dV}\right)_P=\left(\frac{\partial S}{\partial T}\right)_V\left(\frac{dT}{dV}\right)_P+\left(\frac{\partial S}{\partial V}\right)_T$$



Question 1: how does



$$\left(\frac{dS}{dV}\right)_P=\left(\frac{\partial S}{\partial T}\right)_V\left(\frac{dT}{dV}\right)_P+\left(\frac{\partial S}{\partial V}\right)_T$$


become


$$\left(\frac{\partial S}{\partial V}\right)_P=\left(\frac{\partial S}{\partial T}\right)_V\left(\frac{\partial T}{\partial V}\right)_P+\left(\frac{\partial S}{\partial V}\right)_T???$$



Method 2:


Differentiate both side wrt V, holding P const. and use product rule


$$\frac{\partial}{\partial V}\left(dS\right)_P=\frac{\partial}{\partial V}\left(\left(\frac{\partial S}{\partial T}\right)_VdT\right)_P+\frac{\partial}{\partial V}\left(\left(\frac{\partial S}{\partial V}\right)_TdV\right)_P$$


$$\left(\frac{\partial dS}{\partial V}\right)_P=\left(\left(\frac{\partial^2 S}{\partial V \partial T}\right)_V\right)_PdT+\left(\frac{\partial S}{\partial T}\right)_V\left(\frac{\partial dT}{\partial V}\right)_P+\left(\left(\frac{\partial^2 S}{\partial V^2}\right)_T\right)_PdV+\left(\frac{\partial S}{\partial V}\right)_T\left(\frac{\partial dV}{\partial V}\right)_P$$



Question 2: I got so many extra terms, and how to deal with these $$\left(\frac{\partial \text{ d blah}_1}{\partial \text{ blah}_2}\right)_{\text{blah}_3}$$ terms?




Also, on more general grounds:



Question 3: How to partially differentiate a total differential rigorously?


Question 4: Are partial derivatives that differ in only the kept const. term identical in general?




Answer





$$\frac{dS}{dV}=\left(\frac{\partial S}{\partial T}\right)_V\frac{dT}{dV}+\left(\frac{\partial S}{\partial V}\right)_T\frac{dV}{dV}$$




This doesn't make much sense, because is not a well defined expression. The differential $$ \tag{A} dS=\left(\frac{\partial S}{\partial T}\right)_VdT+\left(\frac{\partial S}{\partial V}\right)_TdV$$ is just telling you that if you impose a slight variation on $V$ keeping $T$ constant, $S$ changes accordingly by an amount we denote with $ \left( \frac{\partial S}{\partial V}\right)_T $, so if you ask what is the partial derivative of $S$ with respect to $V$, the answer is tautologically $ \left( \frac{\partial S}{\partial V}\right)_T $ (and of course all the reasonings are the same with $V$ and $T$ exchanged).


You can safely take this (in this context) as the definition of the differential expression (A). In other words, writing (A) is exactly the same as stating that $S$ is a function of the two variables $T$ and $V$.



What is then the meaning of the expression $\left( \frac{\partial S}{\partial V} \right)_P $ ?



It means that you are now considering $T$ itself as a function of $P$ and $V$, call it $\tilde{T}(P,V)$, and effectively asking for the partial derivative of the function $\tilde{S}$ defined by $$ \tilde{S}(P,V) \equiv S(\tilde{T}(P,V),V) $$ with respect to $V$, which is the quantity: $$ \frac{\partial \tilde{S}}{\partial V} (P,V) \equiv \left( \frac{\partial \tilde{S}}{\partial V} \right)_P \equiv \lim_{\epsilon \to 0} \frac{\tilde{S}(P,V+\epsilon) - \tilde{S}(P,V)}{\epsilon}$$ using the usual chain rule you obtain $$ \frac{\partial \tilde{S}}{\partial V} (P,V) = \frac{\partial S}{\partial T}(P,V) \frac{\partial T}{\partial V}(P,V) + \frac{\partial S}{\partial V}(P,V) $$ and this is what is meant by the less rigorous, shorter expression $$ \left( \frac{\partial S}{\partial V} \right)_P = \left( \frac{\partial S}{\partial T} \right)_V \left(\frac{\partial T}{\partial V} \right)_P + \left( \frac{\partial S}{\partial V} \right)_T $$



how does $$\left(\frac{dS}{dV}\right)_P=\left(\frac{\partial S}{\partial T}\right)_V\left(\frac{dT}{dV}\right)_P+\left(\frac{\partial S}{\partial V}\right)_T$$ become $$\left(\frac{\partial S}{\partial V}\right)_P=\left(\frac{\partial S}{\partial T}\right)_V\left(\frac{\partial T}{\partial V}\right)_P+\left(\frac{\partial S}{\partial V}\right)_T???$$




They are the same thing. You have a function $S$ of the two variables $T$ and $V$. What you can do is just differentiating with respect to one or the other, obtaining: $$\frac{\partial S}{\partial T} (T,V) \qquad \text{and} \qquad \frac{\partial S}{\partial V} (T,V) $$ which means (taking the first one for example) the partial derivative of $S$ with respect to $T$, evaluated at the point $T,V$, i.e. the object defined by $$ \tag{B} \frac{\partial S}{\partial T} (T,V) \equiv \lim_{\epsilon \to 0} \frac{S(T+\epsilon,V) - S(T,V)}{\epsilon}$$ Writing this with $d$ instead of $\partial$, in this context, is just notation: $$ \frac{\partial S}{\partial T} (T,V) \equiv \left( \frac{\partial{S}}{\partial T} \right)_V \equiv \frac{d S}{d T} (T,V) $$ the notation with the $(\cdot)_V$ is useful to remark which variable/variables is/are kept constant while deriving (as you can see in (B), where $V$ is not varied).


Also related:




  1. Determine the Dependence of S (Entropy) on V and T




  2. What exactly is the difference between a derivative and a total derivative







$$\left(\frac{\partial dS}{\partial V}\right)_P$$



Please, do not ever write something like this :). A partial derivative is an operation that you can apply to (multi-variable) functions. A differential is not a (multi-variable) function, and its partial derivatives are not defined.


$dS$ means "a little variation of the variable $S$", which can be caused by a corresponding variation of the parameters on which it depends. If you ask what is the variation of $S$ while keeping some other quantity constant, you just divide $dS$ by that quantity (say $dV$) and impose the constraint you want (which is the first method you mentioned). Or for a more rigorous (and more clumsy) approach, you do the partial derivatives of the $\tilde{S}$ defined above. The result is the same.




Are partial derivatives that differ in only the kept const. term identical in general?




No they are not. Consider the following example: let $F$ be a function of the two variables $A$ and $B$, and suppose that $B$ is also a function of other variables, say $A$ and $C$: $$F = F(A,B), \qquad B = B(A,C)$$ Then we have $$ \left( \frac{\partial F}{\partial A} \right)_B \equiv \lim_{\epsilon \to 0} \frac{F(A+\epsilon,B) - F(A,B) }{\epsilon} $$ but with $\left( \frac{\partial F}{\partial A} \right)_C$ we also have to consider that the second argument $B$ of $F$ changes with $A$, then: $$ \left( \frac{\partial F}{\partial A} \right)_C = \left( \frac{\partial F}{\partial A} \right)_B + \left( \frac{\partial F}{\partial B} \right)_A \left( \frac{\partial B}{\partial A} \right)_C $$ The point is that we are now actually deriving another function, call it $\tilde{F}$, defined by $$ \tilde{F}(A,C) \equiv F(A,B(A,C)) $$ Note that the $B$ in $F(A,B)$ and the $B$ in the above definition are very different objects: the former is a number (an independent variable), the latter is a function of the two variables $A$ and $C$.



Note that in all of these reasonings you are always dealing with partial derivatives of functions of many variables (for example the $S(T,V)$ above). What makes it confusing (and I remember being confused myself by this when dealing for the first time with this subject) is the fact that you implicitly, when needed, consider some variables as functions themselves of other variables (just like $T$ that becomes $\tilde{T}(P,V)$ above). This feels pretty natural when you understand what is the exact meaning of the expressions, but can also be confusing at first.


Understanding Stagnation point in pitot fluid

What is stagnation point in fluid mechanics. At the open end of the pitot tube the velocity of the fluid becomes zero.But that should result...