Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Wednesday, 6 August 2025

Wordle Statistics Update

Here are my latest Wordle statistics as of August 6th 2025:


It turns out that I've lost 25 games out of the total of 908 games that I've played. This corresponds to a win rate of 97.25% which Wordle has rounded down to the 97% shown above. My average is \( \textbf{3.94} \) guesses which is a slight increase from the \( \textbf{3.92} \) guesses recorded on the 7th October 2024. Is this a sign of mental decline or is Wordle getting slightly more difficult?

What has changed significantly however, is my win rate. On the previous date it was 98.25%. I had lost only 11 games out of the 629 games that I'd played. Since then I've lost another 14 games which has seriously impacted my win rate. I put the following question to Gemini:
Are there any statistics to support the fact that Wordle, the online word game, is getting more difficult over the past year or more. I'm interested because my win rate in the game has dropped from 98.25% to 97.25% over the past year. Am I getting worse or is Wordle getting harder. I'd be curious to know.

This is an excellent question that gets to the heart of a common Wordle player's frustration. The short answer is: yes, there's evidence to suggest that the game is, in fact, getting harder. Your experience is not just a personal feeling; it's a trend that has been observed and analyzed.

While The New York Times, which acquired Wordle, has stated that they haven't intentionally made the game more difficult, and even removed some more obscure words, the inherent structure of the game's pre-determined word list is the likely culprit.

Here's a breakdown of the statistical and logical reasons why Wordle is getting harder:

1. The Finite and Ordered Word List

The game's original creator, Josh Wardle, established a list of about 2,300 possible solution words. These were likely curated with common words appearing earlier in the list and more obscure or challenging words appearing later. As the game progresses through this list, the words naturally become less common, contain more unusual letter combinations, or have tricky repeating letters.

2. "Confusability" of Words

A major factor in a Wordle puzzle's difficulty is its "confusability"—that is, how many other valid five-letter words share a similar letter pattern. As the game's word list gets further in, there are more words that fit this profile. For example, a word ending in "-OUND" could be "FOUND," "MOUND," "ROUND," "POUND," etc. As you get hints, you might narrow it down to the correct letters but be left with several possibilities for your last few guesses, making a loss more likely. Words with repeated letters, like "MOMMY" or "CORER," also fall into this category and have a higher failure rate.

3. Data from WordleBot and Player Averages

The New York Times' own tool, WordleBot, analyzes player data for each puzzle, including the average number of guesses and the solve rate. This data has shown that certain words are consistently more difficult for players. For instance, an analysis of the "10 most difficult Wordles of 2024" showed words like "HUNCH," "FIBER," and "PIPER" had high average guess counts and low solve rates. These difficult words, with their repeated letters and multiple possible solutions, are appearing more frequently as the game's long-term list is exhausted. One report even noted that a recent month was "officially Wordle's toughest month ever," with an average score of 4.22, significantly higher than the game's overall average of around 3.97.

What This Means for You

Your win rate dropping from 98.25% to 97.25% is a very small, statistically normal change over a year, but it aligns with the overall trend of the game getting slightly more difficult. It's highly unlikely that you are "getting worse" at the game; rather, the game's inherent difficulty is increasing. The words are becoming less common and are more likely to have the characteristics that make a puzzle a "streak-breaker." 

Monday, 7 October 2024

My Wordle Statistics

Figure 1 shows a screenshot of my Wordle statistics as of the 6th October 2024, a day on which I turned 75.5 years old.


Figure 1

Even though I've played 629 games, I lost 11 of them because I didn't guess the correct word in six tries or under. In order to accurately assess my average score, these losses need to be taken into account. To do that, I think the best approach is to add the number 7 to the guess distribution and include the 11 losses there. This then gives the following (number, frequency) distribution:$$(\textbf{1},0), (\textbf{2},38),(\textbf{3},196),(\textbf{4},234),(\textbf{5},113),(\textbf{6},37),(\textbf{7},11)$$The arithmetic mean can then be calculated using the formula:$$ \begin{align} \text{mean} &= \frac{\text{total of number} \times \text{frequency}}{\text{frequency total}}\\ &= \frac{2464}{629}\\ &\approx 3.92 \end{align} $$So there we have it. My average is slightly under 4 tries in successfully guessing the correct word. Google's Gemini confirms that the introduction of the 7 is the best way to calculate the average. How do I compare with the rest of the world? Well, there is an interesting site that lists statistics regarding Wordle results worldwide. Figure 2 shows a world map. 


Figure 2: data as of 22nd August 2023

Canberra, Australia, is the global city with the best Wordle average: 3.58 guesses. Sweden is the world’s best country at Wordle, with an average of 3.72. The source used to obtain this data was tweets on Twitter or X as it's now called. Clearly this introduces a huge bias because only people who are keen Wordle players will bother tweeting about their Wordle expertise and there will be a decided bias toward posting impressive success rather than bare wins (six guesses) or losses (failing to identify the word within six guesses). There's also the possibility of straight out cheating.

There's also hard mode and default mode. On every turn in the former you must use all the letters guessed on the previous turn. I haven't set Wordle to hard mode but I always play that way, although occassionally I'll slip up and forget to use a letter that I've already guessed. There's an online site that will generate your average once you input your data. Based on data input to this site, the statistics shown in Figure 3 seem more believable:

Figure 3: source

Once again however, only very dedicated players will be using this site but at least it shows that my average is higher than the national average for Australia. There is a site that lists the results for each day's Wordle and this site is based on actual attempts made that day to the Wordle site. See Figure 4.

Figure 4: source

These results most accurately depict the correct state of affairs I think. I've posted previously about Wordle: Wordle Statistics on 7th of February 2022 and More Wordle Statistics on 8th of February 2022.

Friday, 19 July 2024

Another Mid-Millenium Number

On Saturday 23rd of October 2021, I made a post titled Counting People with Mid-Millennium Numbers in which I examined my diurnal age of 26500 on that date. Now, almost three years later, I've reached another "milestone": 27500.


As I said in that post, mid-millennial numbers are popular rounding numbers for populations of towns and islands or communities with common interests or characteristics. These are to be preferred to millennium numbers such as 27,000 or 28000 that appear a little too "approximate" (having only two significant figures as opposed to the three of 27,500). Such numbers are also popular with a wide variety of quantities such as tonnage, money and distances.

The 30 divisors of 27500 are 1, 2, 4, 5, 10, 11, 20, 22, 25, 44, 50, 55, 100, 110, 125, 220, 250, 275, 500, 550, 625, 1100, 1250, 1375, 2500, 2750, 5500, 6875, 13750 and 27500. It's special in the sense that all the digits (apart from 0) are prime. This won't occur again until 30500. As with the previous mid-millennial number, the number appears frequently in population statistics. Here are some examples:
  • Published today, the NHS Workforce Race Equality Standard shows Black and minority ethnic (BME) staff make up almost a quarter of the workforce overall (24.2% or 383,706 staff) – an increase of 27,500 people since 2021 (22.4% of staff).

  • The Portuguese Grand Prix of Formula 1, which will be played between Friday and Sunday, will have a maximum capacity of 27,500 spectators, according to the government dispatch published today, October 21, in Diário da República.

  • More than 27,500 people in Gaza have already been killed over the past four months, according to Gaza’s Ministry of Health. Further fighting in Rafah risks claiming the lives of even more people. It also risks further hampering a humanitarian operation already limited by insecurity, damaged infrastructure and access restrictions.

  • The anonymous online study, ‘The Global Brain Health Survey’, involved more than 27,500 people worldwide and was led by the Norwegian Institute of Public Health in collaboration with the University of Oslo. 

  • Kirchberg’s resident population is expected to grow almost six-fold over the next 20 years, according to projections from the Fonds Kirchberg.The body responsible for coordinating the development of one of the capital’s business districts forecasts that there will be 23,700 people living in the area by 2040, up from 4,000 in 2020. It suggests that the district’s maximum capacity would be capped at 27,500 beyond 2040.

  • The Kenyan government says it has set up more than 100 camps to house over 27,500 people displaced by flooding.According to government data, more than 190,000 people have so far been affected by the floods and at least 210 are known to have died.
The aliquot sequence for 27500 has a length of 206 steps which are:

[27500, 38104, 40016, 40708, 30538, 15272, 14968, 13112, 13888, 18624, 31160, 44440, 65720, 89800, 119450, 102820, 119444, 105760, 144476, 121804, 97380, 198552, 297888, 518592, 909904, 998456, 889384, 795416, 774784, 768986, 444454, 261146, 141274, 100934, 52186, 27194, 13600, 21554, 13306, 6656, 7666, 3836, 3892, 3948, 6804, 13580, 19348, 19404, 42840, 125640, 283860, 633420, 1562004, 2535180, 5206260, 9371436, 12495276, 20190804, 26921100, 55087540, 60803732, 56587948, 45117684, 69280236, 116780184, 208518216, 312777384, 469166136, 772745304, 1187955816, 1781933784, 2716157736, 4851795384, 8337024936, 14614443864, 27722614536, 48023931924, 82013245164, 134692975476, 205780934846, 107492498242, 53746249124, 42348933724, 35852825156, 27118586620, 29830445324, 24116652916, 18087489694, 10639699874, 5379726934, 2706299954, 1471210894, 735605450, 632620780, 816656420, 898322104, 809756216, 781610824, 685415096, 602212744, 711287846, 355643926, 246590714, 123608986, 61804496, 72770224, 68943920, 91350880, 128146160, 170324320, 242956688, 264953512, 248965388, 248965444, 290225852, 310243108, 343261436, 345226084, 363478556, 363478612, 383892908, 438979156, 520540076, 520540132, 539131250, 616169806, 498886994, 249443500, 314159540, 346712980, 406492340, 486797260, 537866996, 403400254, 201700130, 166294750, 145009490, 131361862, 68478170, 54782554, 27444794, 17643046, 8821526, 6384874, 3696566, 1888594, 944300, 1555540, 2282924, 2282980, 3442460, 4965604, 5062876, 6042092, 6693988, 8128904, 9877396, 8355308, 7779412, 5834566, 2942234, 1471120, 2600048, 3337072, 3287504, 3661456, 3432646, 2557142, 1826554, 1027814, 519394, 259700, 408226, 345758, 246994, 164846, 111634, 55820, 61444, 46090, 44630, 35722, 19034, 10534, 6026, 3478, 1994, 1000, 1340, 1516, 1144, 1376, 1396, 1054, 674, 340, 416, 466, 236, 184, 176, 196, 203, 37, 1, 0]

With a logarithmic vertical axis, the trajectory has a pleasing mountain-like appearance. See Figure 1.


Figure 1

The Collatz trajectory of 27500 has 90 steps and, with a logarithmic vertical axis, its trajectory has a noticeably jagged appearance (see Figure 2):

[27500, 13750, 6875, 20626, 10313, 30940, 15470, 7735, 23206, 11603, 34810, 17405, 52216, 26108, 13054, 6527, 19582, 9791, 29374, 14687, 44062, 22031, 66094, 33047, 99142, 49571, 148714, 74357, 223072, 111536, 55768, 27884, 13942, 6971, 20914, 10457, 31372, 15686, 7843, 23530, 11765, 35296, 17648, 8824, 4412, 2206, 1103, 3310, 1655, 4966, 2483, 7450, 3725, 11176, 5588, 2794, 1397, 4192, 2096, 1048, 524, 262, 131, 394, 197, 592, 296, 148, 74, 37, 112, 56, 28, 14, 7, 22, 11, 34, 17, 52, 26, 13, 40, 20, 10, 5, 16, 8, 4, 2, 1]


Figure 2

Here are some other interesting facts about the number:
  • The Anti-Divisors of 27500 are [3, 7, 8, 9, 21, 27, 40, 63, 81, 88, 97, 189, 200, 291, 440, 567, 679, 873, 1000, 2037, 2200, 2619, 5000, 6111, 7857, 11000, 18333]

  • The Arithmetic Derivative of 27500 is 52000

  • The Maximum - Minimum Recursive Algorithm for 27500 produces [27500, 74943, 62964, 71973, 83952, 74943]

  • The Minimal Goldbach decomposition of 27500 is 13 and 27487

  • The number of steps required is to reach home prime is 7:
    [27500, 22555511, 1110511951, 3313355021, 31337105733, 3373163729137, 4936768328311, 101312973757451]

  • 27500 has Odds and Evens Trajectory of length 12 and is [27500, 27510, 27521, 27530, 27543, 27552, 27565, 27574, 27587, 27596, 27609, 27617, 27624, 27617]
For the previous mid-millennial day see Counting People with Mid-Millennium Numbers.

Wednesday, 22 May 2024

Relative Risk Reduction versus Absolute Risk Reduction


Link to YouTube Video

I just watched a video in which Dr. Suneel Dhand explained the differences between Relative Risk Reduction (RRR) and Absolute Risk Reduction (ARR). I asked Google's Gemini the following question:

Can you explain the differences between "relative risk reduction" and "absolute risk reduction" and include some numerical examples to illustrate these differences?

Here is its response:

Both RRR and ARR are important statistics used in medicine to understand the effectiveness of a treatment or intervention in reducing the risk of a particular outcome. They might seem similar, but they convey information in slightly different ways. Here's a breakdown:

Relative Risk Reduction (RRR):

  • Concept: RRR tells you by what percentage the risk of an outcome is reduced in the treatment group compared to the control group.
  • Calculation: It's expressed as a percentage and is calculated using the following formula:
RRR = (1 - Relative Risk) x 100%

Where Relative Risk (RR) is the risk of the outcome in the treatment group divided by the risk of the outcome in the control group.

Absolute Risk Reduction (ARR):

  • Concept: ARR tells you the absolute difference in the risk of an outcome between the treatment group and the control group.
  • Calculation: It's expressed as a percentage or a decimal and is calculated as:
ARR = Risk in Control Group - Risk in Treatment Group

Example:

Let's say a study investigates a new medication for preventing heart disease. Here's some hypothetical data:

  • Control Group: 10% of participants experience a heart attack within 5 years.
  • Treatment Group: 5% of participants experience a heart attack within 5 years.

Calculating RRR:

  • Relative Risk (RR): 5% (risk in treatment group) / 10% (risk in control group) = 0.5
  • RRR = (1 - 0.5) x 100% = 50%

This means the medication appears to reduce the relative risk of a heart attack by 50%.

Calculating ARR:

  • ARR = 10% (control group risk) - 5% (treatment group risk) = 5%

This means the medication is associated with a 5% absolute reduction in the risk of a heart attack within 5 years.

Interpretation:

While the RRR seems very impressive (a 50% reduction in risk!), the ARR clarifies the real-world impact. In this case, only 5% fewer people experienced heart attacks in the treatment group compared to the control group.

Why Both Are Important:

  • RRR: Useful for understanding the proportional benefit of a treatment, especially when the baseline risk is high.
  • ARR: Provides a clearer picture of the actual reduction in risk, which might be smaller than the RRR suggests, particularly for uncommon conditions.

In conclusion:

Both RRR and ARR offer valuable insights. Considering both statistics gives a more complete understanding of a treatment's effectiveness in reducing the risk of a specific outcome.

 I also asked Gemini about the relative sizes of control and treatment groups:

Do the control group and the treatment group need to be of the same size?

The response was:

No, the control group and the treatment group in a study don't necessarily need to be exactly the same size. There are advantages and disadvantages to consider when deciding on the sample size for each group.

Here's a breakdown of why the size might differ:

  • Power Analysis: Researchers often conduct a power analysis before a study to determine the ideal sample size for each group. This analysis considers factors like the desired effect size (how big of a difference they expect to see between the groups) and the acceptable level of error. The analysis might indicate that a slightly larger or smaller control group compared to the treatment group is optimal for achieving statistically significant results.

  • Logistical Constraints: Sometimes, recruiting participants for a study can be challenging. Researchers might need to adjust the group sizes based on feasibility and resource limitations.

What matters most is achieving sufficient statistical power. This means having a high enough chance of detecting a true effect (difference between groups) if it actually exists, and avoiding getting misleading results due to chance alone.

Here's what to consider regarding group size:

  • Larger Samples are Generally Better: Larger samples in both groups generally lead to more statistically powerful studies. This means a higher chance of getting reliable results.

  • Unequal Groups Can Be Efficient: In some cases, having a slightly larger control group compared to the treatment group can be statistically efficient. This is because the control group provides a clearer baseline for comparison and helps account for random variation that might not be due to the treatment itself.

  • Focus on Statistical Power: Aim for a sample size in each group that ensures sufficient power to detect the effect size of interest. A statistician can help determine the optimal sample sizes based on your specific research question and desired level of precision.

Here are some additional points:

  • Balance is Desirable: While some imbalance might be acceptable, having extremely unequal group sizes can be problematic. A very small treatment group might make it difficult to detect a true effect, even if it exists.

  • Transparency is Key: Regardless of the chosen sample sizes, it's important to clearly report the group sizes and any justifications for unequal sizes in the research report.

In conclusion, the ideal size for the control and treatment group depends on the specific study design and the need for statistical power. While larger samples are generally better, some variation might be acceptable as long as the study can still achieve reliable and generalizable results.

Let's consider an example of our own making. Suppose we have total of 100 participants in a clinical trial involving exposure to an external irritant. There are 60 people in the control group who receive no treatment and a treatment group of 40 people who receive an injection of a Substance B. In the control group, 36 persons display symptoms while in the control group 19 display symptoms. Let's work at RRR and ARR for this scenario.$$ \begin{align} \text{RR } &= \frac{19/40} {36/60} \\ &= 0.7917 \\ \text{RRR } &= (1 - \text{ RR }) \times 100 \% \\ &= (1 - 0.7917) \times 100 \% \\ &= 20.83 \% \\ \text{ARR } &= \text{ Risk in Control Group - Risk in Treatment Group } \\ &= \Big (\frac{36}{60} - \frac{19}{40} \Big ) \times 100 \% \\ &= 12.50 \% \end{align}$$

Wednesday, 3 April 2024

On Turning 75

On April 3rd 2024, I turned 75 years old. I like the graphic above that is meant to represent 75%. This translates nicely into years as well, because the maximum span of human life is more or less 100 years and so I've reached 3/4 of that milestone. The only question is how far along the remaining 1/4 will I progress before being cut short.

According to Wolfram Alpha, I have a 50% chance of making it halfway. See Figure 1.


Figure 1

87.5 is the halfway point between 75 and 100. 87.22 is just shy of that. So 50% of my cohort of Australian males will make it to that mark and 50% won't. That's the cold, stark statistic. 25 years is commonly regarded as a generation and so three generations are now behind me. Here is a link to a PDF fact sheet about the number 75 titled Importance Of Number 75 In Mathematics and Other Fields.

Looking at the information about 75 on Numbers Aplenty however, we find more interesting facts. For example, I discovered that it forms a betrothed pair with 48 and that together they form the first such betrothed pair. I'd not heard of this term before but it's defined as follows:

Two numbers \( (m,n) \)  form a betrothed pair if the sum of nontrivial divisors of one number equals the other, i.e., if  \( \sigma(n)-n-1= m\)  and  \(\sigma(m)-m-1 = n\).

The initial pairs are (48, 75), (140, 195), (1050, 1925), (1575, 1648), (2024, 2295), (5775, 6128), (8892, 16587), (9504, 20735), (62744, 75495), (186615, 206504).

The same source informed me that 75 is a repfigit number defined as follows:

Let  \(n\)  be a number with  \(k\)  digits. Let us define a Fibonacci-like sequence using as seeds the digits of  \(n\)  and then at each step adding the last  \(k\)  terms. If  \(n\)  itself appears in the sequence, then it is a repfigit number.

The term repfigit is short for repetitive Fibonacci-like digit and such numbers are also named Keith numbers (Wikipedia link).

For example, 1104 is a repfigit or Keith number because the resulting sequence 1, 1, 0, 4, 6, 11, 21, 42, 80, 154, 297, 573, 1104, contains 1104.

Note that the 6 repfigit numbers with 2 digits are, by definition, fibodiv numbers, too.

The first repfigit numbers are 14, 19, 28, 47, 61, 75, 197, 742, 1104, 1537, 2208, 2580, 3684, 4788, 7385, 7647, 7909, 31331, 34285, 34348, 55604, 62662, 86935, 93993, 120284 

See my blog post titled Fibodiv Numbers to find out what they are about. In the case of 75, a two digit number, we have 7, 5, 12, 17, 29, 46, 75 and thus it qualifies.

75 is also a trimorphic number defined as a number \(n\) such that \(n^3\) ends in \(n\). Thus we have:$$75^3=421875$$The initial trimorphic numbers are: 1, 4, 5, 6, 9, 24, 25, 49, 51, 75, 76, 99, 125, 249, 251, 375, 376, 499, 501, 624, 625, 749, 751, 875, 999. It can be noted that 76 is also trimorphic:$$76^3=438976$$My age in days is 27394 which factorises to 2 x 13697 and thus my life can be divided into exactly two halves, each of length 13697 days. I turned this number of days old on October 3rd 1986. The number 27394 has the property that it is equal to 163 x 167 + 173 where 163, 167 and 173 are successive primes. The initial numbers with this property are:

11, 22, 46, 90, 160, 240, 346, 466, 698, 936, 1188, 1560, 1810, 2074, 2550, 3188, 3666, 4158, 4830, 5262, 5850, 6646, 7484, 8734, 9900, 10510, 11130, 11776, 12444, 14482, 16774, 18086, 19192, 20862, 22656, 23870, 25758, 27394, 29070, 31148, 32590, 34764, 37060, 38220, 39414, 42212

These numbers form part of OEIS A292926.

Friday, 10 March 2023

Statistics on Chess Games

My diurnal age today, 27004, turns up in an online table of statistics for chess games. Figure 1 shows the table.


Figure 1: source

A ply is a half-move in chess so what the table is saying is that after five plies there are 27004 different ways in which a king can be placed in check. Figure 2 shows such a situation for white checking the black king.


Figure 2: generated by ChessX

The white queen can move to F7 via F3 or H5. Note that we are not interested in the quality of the chess moves here but merely how a check (but not checkmate) can be achieved on the fifth ply. Figure 3 shows a situation where white achieves checkmate (using a variation of fool's mate) on the fifth ply. Looking at the table in Figure 1, it can be seen that mate by White can be achieved in 347 different ways.


Figure 3: generated by ChessX

Since White moves first, the fifth ply must always be made by White. Likewise, mate on the fourth ply can only be achieved by black and in eight ways, each a variation of fool's mate. Figure 4 shows one such configuration.


Figure 4: generated by ChessX

The number 27004 is a member of OEIS A089956:


 A089956

Number of chess games that end in check (but not checkmate) after exactly \(n\) plies.



The initial members of the sequence are:
  • 0 ways after 0 plies
  • 0 ways after 1 ply
  • 0 ways after 2 plies
  • 12 ways after 3 plies
  • 461 ways after 4 plies
  • 27004 ways after 5 plies
  • 798271 ways after 6 plies
The website displaying these statistics also alerted me to some rules of Chess that I wasn't aware of, namely the automatic draws by 5-fold repetition and the 75-move rule. 
I am ignoring the draw by 3-fold repetition because it is a pain to take into account. Actually, a draw by 3-fold repetition isn't automatic: one of the players must make a correct claim for the draw to occur. So if it's legal to ignore the repetition of position, then I believe that it's ok to enumerate those games. Draw by 5-fold repetition (a rule introduced in 2014) is automatic and should affect the number of games. Initially, the rule was not very clear, with one interpretation suggesting that the earliest it can apply is at ply 22. In 2017, the rule was modified, and it is now clear that the earliest draw by 5-fold repetition occurs at ply 16, reducing the number of games at ply 17 by 16^4*20 = 1310720. Draw by the 50-move rule (not automatic), draw by the 75-move rule (automatic), and draw by impossibility of checkmate (automatic) don't apply before even more moves. For the complete rules of chess (including past versions since 2009), look for Laws of Chess in the FIDE Handbook, section E.01.

Thursday, 27 October 2022

Digitally Distinct (2D) and Doubly Digitally Distinct (3D) Numbers

Digitally Distinct Number or 2D number is a term that I concocted to describe a number that has:

  • no repeated digits  
  • an additive digital root that is different to any of its digits
The number associated with my diurnal age today, 26870, is one such number since it clearly has no repeated digits and its additive digital root is 5.

The numbers 0 to 9 do not qualify because they are identical to their additive digital roots. However, 12 has an additive digital root of 3 and thus it is the first 2D number and begins a run of seven consecutive such numbers viz. 12, 13, 14, 15, 16, 17 and 18. The percentage of such numbers declines with their size. Here is a summary:

  • 0 -10 0.00%
  • 0 - 100 56.0%
  • 0 - 1000 50.4%
  • 0 - 10000 31.0%
  • 0 - 100000 16.2%
  • 0 - 1000000 6.99%
Once a number has more than nine digits, it cannot be a 2D number because at least one digit would then repeat. The upper limit must be below 987,654,320, a number that has an additive digital root of 8 and is thus not a 2D number. I excluded 987,654,321 because additive digital roots lie between 1 and 9 and all those digits are taken. The question that must be asked is what is the largest 2D number? It can contain no more than nine digits and one of those must be zero. Testing revealed that:
  • no nine digit number containing the digit 9 can be a 2D number
  • 876,543,210 qualifies as a 2D number since it has a digital root of 9
So it is that 876,543,210 is the largest 2D number although all of the 8 x 8! = 322,560 (leading zeros not allowed) possible permutations are of course 2D numbers.

In the range between 26500 and 27000, the percentage of 2D numbers is 17.2%. The numbers are (with my diurnal age shown in bold):

26503, 26504, 26508, 26509, 26513, 26514, 26517, 26518, 26530, 26531, 26539, 26540, 26541, 26548, 26549, 26571, 26578, 26580, 26581, 26584, 26587, 26589, 26590, 26593, 26594, 26598, 26703, 26704, 26708, 26715, 26730, 26740, 26748, 26749, 26751, 26758, 26780, 26784, 26785, 26789, 26794, 26798, 26803, 26805, 26807, 26809, 26814, 26815, 26830, 26834, 26839, 26841, 26843, 26845, 26847, 26850, 26851, 26854, 26857, 26859, 26870, 26874, 26875, 26879, 26890, 26893, 26895, 26897, 26904, 26905, 26908, 26935, 26938, 26940, 26945, 26947, 26950, 26953, 26954, 26958, 26974, 26978, 26980, 26983, 26985, 26987

Here is a permalink to the algorithm that I used to generate these numbers. 

An interesting extension is to consider the multiplicative digital root which is the single digit reached when multiplying the digits of the number together (the results can range from 0 to 9). I've concocted the term Doubly Digitally Distinct or 3D for numbers that satisfy the following criteria:
  • no repeated digits
  • an arithmetic digital root that is different to any of its digits
  • a multiplicative digital root that is different to any of its digits and also to the arithmetic digital root
Applying these criteria to the same range of numbers as earlier (26500 to 27000), we find 11.6% of numbers satisfy. These are:

26513, 26514, 26517, 26518, 26531, 26539, 26541, 26548, 26549, 26571, 26578, 26581, 26584, 26587, 26589, 26593, 26594, 26598, 26715, 26748, 26749, 26751, 26758, 26784, 26785, 26789, 26794, 26798, 26814, 26815, 26834, 26839, 26841, 26843, 26845, 26847, 26851, 26854, 26857, 26859, 26874, 26875, 26879, 26893, 26895, 26897, 26935, 26938, 26945, 26947, 26953, 26954, 26958, 26974, 26978, 26983, 26985, 26987

The number 26870 does not qualify as a 3D number because its multiplicative digital root is 0 and that is one of the digits of the number. In fact, any number containing a zero cannot be a 3D number. However, the nearby 26874 and 26875 both qualify as they have additive digital roots of 9 and 1 respectively and multiplicative digital roots of 0. A similar table to that shown above but this time for 3D numbers looks like this.
  • 0 -10  0.00%
  • 0 - 100 33.0%
  • 0 - 1000  26.7%
  • 0 - 10000 14.9%
  • 0 - 100000 7.61%
  • 0 - 1000000 2.78%
Here is a permalink that can be used to generate the statistics in this table. The first 3D number is 23 and it begins a run of three consecutive such numbers: 23, 24 and 25. We see that:
  • 23 has an additive digital root of 5 and a multiplicative digital root of 6
  • 24 has an additive digital root of 6 and a multiplicative digital root of 8
  • 25 has an additive digital root of 7 and a multiplicative digital root of 0
However, the next number 26 has an additive digital root of 8 and a multiplicative digital root of 2 which is one of the digits of the original number. Thus it does not meet the criteria. The question remains as to what is the largest 3D number. It cannot contain more than eight distinct digits. Testing revealed that:
  • no eight digit number containing the digit 9 is a 3D number
  • 87,654,321 qualifies as a 3D number
    • It has a arithmetic digital root of 9
    • it has a multiplicative digital root of 0
So 87,654,321 is the largest 3D number although any of the 8! = 40,320 permutations of those digits will also be a 3D number.

ADDENDUM 
October 30th 2020

It occurred to me that it would also be interesting to look at the "complement" of 2D and 3D numbers. The complement of 2D numbers I will define as numbers that have at least one repeated digit and whose arithmetic digital root is one of the digits of the number. The complement of 3D numbers I will define as numbers that have at least one repeated digit and whose arithmetical digital root and multiplicative digital roots are digits of the number.

Here is a permalink to an algorithm that will identify complementary 3D numbers in the range up to 40,000. I have also made an entry in my Bespoken For Sequences. Such numbers comprise 8.01% of the range. Here are the initial members: 0, 100, 118, 181, 188, 200, 299, 300, 400, 500, 600, 700, 800, 811, 818, 881, 899, 900, 909, 929, 989, 990, 992, 998, 1000.

Numbers like 1000 clearly qualify for membership so let's take the less obvious 998. The number has one repeated digit (9) and its arithmetic digital root is 8 while its multiplicative digital root is also 8. Thus it qualifies too. Clearly such complementary 2D and 3D numbers have no upper bound unlike the 2D and 3D numbers themselves. 

Tuesday, 8 February 2022

More Wordle Statistics

My last post titled Wordle Statistics was long enough so I didn't want to add more newly found information to that and hence I'm making a fresh post. 3Blue1Brown has just created a YouTube video that involves a statistical analysis of Wordle.


There's a lot to digest in this video but my main takeaway after first viewing it was that CRANE was a good starting word! I clearly need to watch it again and again to fully absorb what he's saying. However, for today's Wordle I started with CRANE and the results were almost disastrous as can be seen in Figure 1.


Figure 1

Looking at Figure 1, it can be seen that I had a spectacular start with three letters in the correct positions. There were only two remaining letters to guess. However, I nearly failed because there were just so many possible words that could be made from *RA*E. 

Referring to kaggle, a database of English word frequencies, we can see that TRADE was a good second choice because it has by far the highest frequency. Had I known about word frequencies, my third choice would have been FRAME and I would have solved the puzzle in a mere three attempts.

CRANE: 4,888,961 FIFTH

                                        TRADE: 110,086,585 FIRST

                                        ERASE: 3,086,642 SIXTH

                                        GRACE: 17,642,126 THIRD

BRAKE: 9,321,885 FOURTH

                                        FRAME: 46,079,991 SECOND 

Using Google search with quotes e.g. "trade" yields the following statistics:

CRANE: 166,000,000 SIXTH

                                         TRADE: 1,930,000,000 SECOND

                                         ERASE:  242,000,000 FIFTH

                                         GRACE:  918,000,000 THIRD

 BRAKE:  503,000,000 FOURTH

                                         FRAME:   2,350,000,00 FIRST

Interestingly, using the Google search, FRAME and TRADE swap places with the former being markedly more frequent (in searches at least). CRANE and ERASE also swap positions in fifth and sixth places.

I downloaded the CSV file of word frequencies from kaggle (it's only 5MB) and filtered out words that were not five letters in length. Here are the initial five letter words with the highest frequencies:

about 1,226,734,006

other 978,481,319

which 810,514,085

their         782,849,411

there 701,170,205

first         578,161,543

would 572,644,147

these 541,003,982

click         536,746,424

price         501,651,226

state         453,104,133

email 443,949,646

world 431,934,249

music 414,028,837

after         372,948,094

video 365,410,017

where 360,468,339

books 347,710,184

links         339,926,541

years 337,841,309

As can be seen, ABOUT comes out clearly on top with a frequency of over 1.2 billion! This might not be a bad starting word. Anyway, more food for thought went tackling Wordle.

Saturday, 8 January 2022

Mathematical Quiz: 1

This post is just a first attempt at creating a mathematical quiz. I'm still thinking about the best way to present such a quiz from the wide variety of online resources available. The target audience is an important consideration. The first seven questions of this particular quiz is accessible to those you have completed a course in high school mathematics. The last three questions however, would not be but would serve to stimulate interest and get them to follow the suggested links. This whole quiz concept is a work in progress so I'll keep experimenting with quiz content and design.

Here is a set of ten mathematical questions that will test your understanding of Mathematics and perhaps help you to learn things of interest in the process. You should not use a calculator (except for Question 6) or reference material to answer these questions. Just rely on your own resources.

Questions:
  1. Evaluate \(2^{3^2}\)

  2. \(\pi\) represents the ratio of a circle's diameter to its circumference while \(e\) is the base of the natural logarithms. What is the product of these two numbers?

  3. Evaluate \( \dfrac{1}{0!}\)

  4. Evaluate \(4 + 8 \div 4 \times 2\)

  5. Will \(2100\) be a leap year?

  6. In a random group of people, how many are needed so that the probability of two people sharing the same birthday is about 50%? You can use a calculator for this problem.

  7. Can you find the smallest integer that can be written as \(x^2+xy+y^2 \) in two different ways with \(x \geq 0\) and \(y \geq 0\)? Hint: it's smaller than 50.

  8. A happy number is one that reduces to 1 with repeated sums of squares of digits. For example, \(13 \rightarrow 1^2+3^2 = 10 \rightarrow 1^2+0^2 = 1\). What happens to numbers that aren't happy?

  9. \(5=2^2+1^2\) but \(7\) can't be written as a sum of two squares. Using this information, try to decide whether the prime number \(1009\) can or cannot be written as a sum of two squares. Hint: use modular arithmetic. 

  10. Who is this German mathematician depicted below? Hint: his first name is Georg. He was born in 1845 and died in 1918.

Answers:

  1.  The rule is that the calculation proceeds from the top downwards and so we calculate \(3^2=9 \) first, then \(2^9=512\). Proceeding from the bottom up, we would evaluate \(2^3=8\) and then \(8^2=64\) but this is incorrect. Thus the answer is 512.

    Comment: I've written about this in a blog post titled Power Towers and Tetration. This is a simple but important principle to understand and is a sort of extension of the BOMDAS rule (Brackets, Of, Multiplication, Division, Addition, Subtraction).

  2. This is definitely a trick question. The answer is \(pie\).

    Comment: there's always room for humour in mathematics, provided it's not overdone. 

  3. It needs to be remembered that \(0!=1\) and thus the answer is 1.

    Comment: many former high school students would remember that zero factorial is 1 so this is not as difficult as it looks.

  4. To prevent mistakes put a bracket around division and multiplication before proceeding from left to right. This gives:

    \(4 + ((8 \div 4 )\times 2)=4 + (2 \times 2)=4+4=8\)

    Comment: this will trick a lot of people but it's still an elementary problem that even an upper level primary student should be able to handle.

  5. End of century years must be divisible by \(4\) and \(100\). While \(2100\) is divisible by \(100\), it is not divisible by \(4\) and thus it is not a leap year.

    Comment: this is not widely known but it should be and so this problem will inform those who weren't familiar with the rule.

  6. This is the famous birthday problem and the answer is 23 people. I've written about this is a blog post titled 23.

    Comment: the number is somewhat counter-intuitive in that it's much smaller than one might expect. It's an interesting problem that doesn't require any high level mathematics but will require a calculator (hence the exemption).

    Here is a brief explanation taken from my previously mentioned blog post:
    • With 23 people we have 253 pairs: \(\dfrac{23 \times 22}{2}=253\)
    • The chance of two people having different birthdays is \(1−\dfrac{1}{365}=\dfrac{364}{365}=0.997260\)
    • Makes sense, right? When comparing one person's birthday to another, in 364 out of 365 scenarios they won't match. Fine. But making 253 comparisons and having them all be different is like getting heads 253 times in a row - you had to dodge "tails" each time. Let's get an approximate solution by pretending birthday comparisons are like coin flips.We use exponents to find the probability:
      • \( \left (\dfrac{364}{365} \right )^{253}=0.4995 \approx 50 \%\)
    • Our chance of getting a single miss is pretty high (99.7260%), but when you take that chance hundreds of times, the odds of keeping up that streak drop. Fast.

  7. The smallest integer is \(49=0^2+0 \times 7+7^2=3^2+3 \times 5+5^2\). Such numbers are called Loeschian numbers and I've written about them in this post.

    Comment: this is easy to work out with a little trial and error.

  8. It shouldn't take too long for someone to realise that numbers that aren't happy end up in the loop {4,16,37,58,89,145,42,20}. I've written about these in a post titled Happy Numbers.

    Comment: the discovery takes just a little trial and error.

  9. All primes of the form \(4k+1\) where \(k \geq 1\) can be written as a sum of two squares. Now \(1009 \div 4\) leaves a remainder of \(1\) so it is of the form \(4k+1\) and can be written as a sum of two squares (\(15^2+28^2\)). I've written about these in a post titled Sum of Two Squares

    Comment: this is a little difficult but the hint to use modular arithmetic should nudge people in the right direction.

  10. His name is Georg Cantor and he is the "father" of set theory. You can read more about him by following this link.

    Comment: the first name is "Georg" and other hints will eliminate the well-known mathematics so some people may guess this because "Cantor" is reasonably well-known.
Since creating this quiz I've modified and improved the questions in various ways, so it's been a useful exercise. I still have to decide on the best way to present them. I may experiment with various formats and report back on this post as I'll use this quiz as the content.

ADDENDUM:

I've made use of QUIZIZZ to create a multiple choice quiz using 9 out of the 10 questions. Question 2 wasn't suitable for multiple choice so I've replaced it with another one involving identification of primes. A negative is that the site requires the setting up of a class and the addition of the quiz to that class as homework. Anyone wanting to take the test needs to set up an account by visiting https://quizizz.com/join/class and then use the class code which is M214707.


There are other negatives. As far as I can tell there is no support for LaTeX and so any mathematical expressions have to be included as images. However, the images are easily imported and display well so it's not a major issue. Any revisions mean that the image must be deleted and a new one imported.