How piefed currently ranks
I took a look into post.post_ranking which is an implementation of what has been described by Amir Salihefendic in 2015.
if post_date is None:
post_date = datetime.utcnow()
if score is None:
score = 1
order = math.log(max(abs(score), 1), 10)
sign = 1 if score > 0 else -1 if score < 0 else 0
seconds = self.epoch_seconds(post_date) - 1685766018
return round(sign * order + seconds / 45000, 7)
It looks esotheric to me.
- The log(… score) is quickly explained by the recursion of higher visibility through higher scores. The higher the score, the more views and more votes. That’s fine with me. A lot depends on the first votes, I can’t argue with that. But why they picked the base of 10 is left unexplained.
- The value 1685766018 refers to the unix date
2023-01-03T14:07:42. It looks like a magic/random date to me. However that date seems to be unimportant later on. - The divisor of 45000 seconds is equivalent to 12.5 hours, 1 half-day. Through dividing the post age by 12.5 hours the age is basically converted from seconds to half-days. Again, unexplained why 12.5 h.
- In the end, we add up log(score, 10) + age in half days. I don’t get why they add them up.
Say, we choose a different log base like 2 instead of 10. How would the age measurement need to be adapted in order to get the same post_ranking results? All I know is it is unimportant. We just have to adapt the 45000 s to be 13546 s.
If we’d change the magic date we would need to adapt the arbitrary 45000 seconds again. (Wouldn’t change the rank, I suspect.)
Different approaches / statements
Here, something slightly different is claimed to be the reddit formula: Here, the sign decides if time is good or bad, not the log score.
score = log_10(|score|) + sign(score) * seconds / 45000
Lobsters has this: (I removed a hotness-bonus, a comment score and several plus-ones.)
score = sign(score) * log(|score| + seconds / 81000)
They call the 81000 s (22.5 hours) the “hotness window”. That article features an interesting animation of three kinds of posts: viral hit, sleeper hit, steady performer.
Evan Miller explored the reddit formula and boiled it down to:
ln( score ) + factor * seconds
Seeing similarites to Bayesian beta distribution in a poisson process, Miller concludes:
I realize that proposing any change to how Reddit works is one of the Internet’s most dangerous games, so I hesitate to beat the drum in favor of MillerSort™. But I believe that expected-utility theory and a simple random-reload model can help explain why the Reddit formula has been so effective in the past, and shine a light on aspects that might be improved. In particular, the Reddit formula should probably take into account the percent of votes that are positive, rather than just taking the difference between positive and negative votes.
By the way, the original seems to be over 12 years old: https://github.com/reddit-archive/reddit/commit/50d35de04b928836b7ee955c8a26f197e24ab01e That commit explains the diversity of where to put the sign function.
Say, I’d suggest a different, more physical ranking function. How would we test it? Does somebody have statistics?
Edit: Correction concerning the choice of the log base. Precision on point 4. Guess about having to change 45000 for a different magic date.
Edit 2: Moved my comment beneath a new headline in my OP.
Edit3: Added more different approaches / statements. Consistent naming.
Edit4: Fixed an error in Lobsters’ function.
