What the Gini index is and why you need an adjusted one
Say you are looking at these numbers
| Metric | Value |
|---|---|
| Posts | 50 |
| Authors | 5 |
| Views | 1,000,000 |
| Metric | Value |
|---|---|
| Views per post | 20,000 |
| Views per author | 200,000 |
Can you draw conclusions from this? Probably, but they will not be objective.
Per author the numbers look like this:
| Author | Posts | Views | Share |
|---|---|---|---|
| @alice | 10 | 990,000 | 99.00% |
| @bob | 10 | 2,500 | 0.25% |
| @carol | 10 | 2,500 | 0.25% |
| @dave | 10 | 2,500 | 0.25% |
| @erin | 10 | 2,500 | 0.25% |
For a saner picture we can look at percentiles
| Percentile | Views per author |
|---|---|
| 10 | 2,500 |
| 25 | 2,500 |
| 50 (median) | 2,500 |
| 75 | 2,500 |
| mean | 200,000 |
Now we know for sure that 75% of authors get 2,500 views or fewer. But what if we want to express this as a single number? That is where the Gini index helps.
The Gini index is a measure of distribution inequality. 0 means every author collects the same, 1 means one author takes everything.
| Symbol | Meaning |
|---|---|
| number of authors in the community | |
| views (or posts) of the author in -th place, with all authors sorted ascending: | |
| the author's position in that sorted list, from 1 to | |
| total views of the community |
As an example, take this community: The Startup.
The standard formula was meant for large datasets, so when we apply it to X communities with few authors the numbers come out wrong. For example:
Let , so . In the sum only the last term is non-zero, , and total views are 100:
Even with a total skew toward one author, the index comes out at 0.8.
The maximum of the index over authors is , not 1.
| Authors | Maximum Gini | Error |
|---|---|---|
| 2 | 0.500 | 50.0% |
| 3 | 0.667 | 33.3% |
| 5 | 0.800 | 20.0% |
| 10 | 0.900 | 10.0% |
| 20 | 0.950 | 5.0% |
| 50 | 0.980 | 2.0% |
| 100 | 0.990 | 1.0% |
So for fewer than 50 authors we use the adjusted Gini index:
On the same example: .
Use x-community.top for proper community analysis.