# Are the red and the blue walkers the same kind of walker? Two crowds start at home. Each red walker takes a step of length $s_A$ in a random direction every tick, each blue walker a step of length $s_B$, and each step may carry a small drift $v$ added to its $x$-component. After $n$ ticks you are handed the two clouds and asked whether they came from the same population. That is a hypothesis test, and the point of this page is that *which* test you run decides what you can see. <div class="applet" data-applet="two-walkers"></div> ## What the two clouds look like in theory A step is $(s\cos\theta + v,\; s\sin\theta)$ with $\theta$ uniform, so after $n$ steps $ \mathbb E[x_n] = v\,n, \qquad \operatorname{Var}(x_n) = \frac{n s^2}{2}, \qquad \mathbb E[R_n] \approx \frac{s}{2}\sqrt{\pi n}\ \ (v = 0), $ and by the central limit theorem $x_n$ is close to normal. So the three things that can distinguish the populations live in different places: **drift** moves the mean of $x$; **step size** leaves the mean of $x$ alone and changes the variance of $x$, and with it the mean distance $R$ from home. ## The three tests **Welch's $t$ on the mean of $x$.** With group summaries $(\bar x_A, s_A^2, n_A)$ and $(\bar x_B, s_B^2, n_B)$, $ t = \frac{\bar x_A - \bar x_B}{\sqrt{s_A^2/n_A + s_B^2/n_B}}, $ referred to a $t$ distribution with the Welch–Satterthwaite degrees of freedom. It does not assume equal variances, which matters here, since unequal variance is one of the things we are looking for. **Welch's $t$ on the mean of $R$.** The same statistic applied to the distances from home. **The $F$-test on the variance of $x$.** $F = s_A^2 / s_B^2$ on $(n_A - 1,\, n_B - 1)$ degrees of freedom, two-sided. The $F$-test is famously sensitive to non-normality, which is why it is a reasonable choice here and often a poor one elsewhere: $x_n$ is a sum of $n$ independent steps and is very nearly normal. The two-sided $p$-values come from the regularized incomplete beta function, $p_t = I_{\nu/(\nu + t^2)}\!\left(\tfrac{\nu}{2}, \tfrac12\right)$ and $P(F \le f) = I_{d_1 f/(d_1 f + d_2)}\!\left(\tfrac{d_1}{2}, \tfrac{d_2}{2}\right)$. ## What to check **Same population.** Press *same population* and run. The three $p$-values wander, and every so often one of them dips under $\alpha$: that is a Type I error, and it is supposed to happen. Press *repeat the experiment*: two hundred fresh runs, and each test rejects in about 5% of them at $\alpha = 0.05$. A test that never rejected a true null would be a test that never rejected anything. **B steps further** ($s_B = 1.3$). The mean-of-$x$ test sees nothing, because both means are zero; the $F$-test and the mean-of-$R$ test both reject. Now the surprise: run to 400 ticks instead of 100 and repeat the experiment. The powers do not move. The difference in $\bar R$ grows like $(s_B - s_A)\sqrt n$, but so does the spread of $R$, so the $t$-statistic is independent of $n$; the variance ratio is $s_B^2/s_A^2$ at every $n$. For a step-size difference, more ticks buy nothing. More *walkers* do. **B drifts** ($v_B = 0.05$). Now the mean-of-$x$ test is the one that works, and its power grows with the tick count: the mean separation is $v n$ while the standard error is $\propto s\sqrt{n}$, so $t \propto \sqrt n$. Watch its $p$-value fall steadily in the lower panel while the $F$-test's $p$-value stays flat and uninformative. The mean-of-$R$ test catches up eventually, since a drifting cloud is also a far-from-home cloud. **Both.** All three reject, for different reasons. Read the readout: each test is reporting a different fact about the two populations. **$\alpha$.** Set it to 0.01 and repeat the *same population* experiment. Fewer false alarms, and in the *B steps further* case, less power. That is the trade, and nothing about the walk changes it. **Two hundred is not many.** The rejection rates the repeat button reports are themselves estimates from a finite sample: at 200 repetitions a true 5% rate comes out anywhere from about 2% to 8%. Slide the repetitions up to 1000 if you want the second digit. --- The walkers are those of Wilensky's NetLogo Random Walk 360 model, rebuilt on [[Random Walks]]; the two-sample tests and the repeated-experiment machinery are the additions here. Companion note on the fenced and goal-seeking versions: [[Random Walk 360 (NetLogo)]].