Neo4j Performance adventures for petabyte-scale datasets

I recently proposed a new research project idea: let’s take all of GitHub (or <insert your preferred VCS host>) and create a multi-language (even partially language-agnostic) concrete syntax tree of all the code so that we can do some otherwise impossibly difficult further research and answer incredibly complex questions.

This project is named World Syntax Tree, or WST in short.

Originally I started the project using MongoDB and storing references between nodes as ObjectIDs, but I quickly realized that a tabular format was not performant enough to be able to effectively represent a true tree.

So instead I switched over the whole project to the first and foremost graph database I came across: Neo4j.

As I quickly learned the new database paradigm I also quickly learned that there are a lot of problems between me and inserting literally hundreds of terabytes of data into a single graph…

With Factorio 1.1 a simple yet game-changing feature was introduced: Train Stop Limits.

Normally I always stick to LTN or TSM, or , but now I'm feeling comfortable with a completely vanilla logisitic train network.

With just a few combinators we can emulate the logistic network with trains almost exactly. (Still no easy support for cargo-agnostic train deliveries though.)

So what does it take to have a logistic train network in unmodded Factorio?

Well turns out we can break down the steps fairly easily…

CloudFlare just opened up it’s Web Analytics to everyone.

Their proposal is that their solution is far more privacy friendly compared to other conventional analytics like Google, which totally makes sense, but how does it compare to my current choice: Matomo?

Drop X THX Panda: Perfectly clear

Updates: 3 Months

After many months of anticipation and multiple delays in production, I finally have on my head a pair of “the world’s highest fidelity wireless headphones.”

And I think I will agree with their claim - although my word may not mean as much as some professional reviewers out there with $500+ pairs of headphones, I find that these headphones have clarity far beyond anything I’ve tried and reproduce sound so accurately it’s somewhat spooky.

A lot of people will probably disagree with me, I am no expert in audio, after all this is my first planar headphone experience ever, so instead of focusing so much on their pure audio quality, I’ll talk about my specific experience with them.

Dear Keybase, I am very scared.

Read Keybase’s blog post here.

If all you read is the intro text, take this quote:

A bought-out company can never be trusted more than the parent company.

And unfortunately, Keybase, a company which I originally held in extremely high regard, just got bought by one which I personally have strong negative prejudice about.

And thus marks the downfall of Keybase’s trust factor.

If this were the other way around: If Keybase acquired Zoom: I would be ecstatic, because I truly think Keybase has the ability to create a great end-to-end complete platform for all kinds of communication, but not quite the best possible platform, as you’ll see lower down in this post.

Though luckily, not all is completely lost: