Migrating To Mongo
I have an application which contains data structures that did not fit nicely in an SQL database due to the nested structure of the data, so instead I opted to only store the metadata in a PostgreSQL database and the actual content as a plain JSON file on disk. It was however time to move them into a MongoDB database.
Before we dive into the details, let me give you some context as to why this move was required. The thing is that storing the content in a file works fine as long as you don't want to search through the content, but of course the point came where I had to filter/group the files based on some content inside of it. Since the content only existed on disk, this means reading all files, parsing them and then applying the filter. I did not test it, but I assumed that such a search would only be feasible for a small quantity of files.
I have always wanted to play around and get some experience with MongoDB, so it seemed like this was the perfect opportunity to do so. Initially I didn't expect too many issues since MongoDB is a document store and it handles data in a JSON structure, which my data already was. So it was just a matter of loading all data in the database, and loading it from there from now on. Well, there are a couple of caveats. First of all, MongoDB uses a type of UUID, where all of my id's where simple integers. This means I had to have some mapping from my PostgreSQL id to the MongoDB id. Nothing too complicated, though.
Since this is an application that is actually used by other people as well, I wanted to avoid a huge catastrophe with the new version. Therefore, I opted for a multi-phase approach where I kept the original metadata in the PostgreSQL database and the content as a file on disk, but I started to also store it in MongoDB. Reading the data would always be from MongoDB. This allows me to verify that data is correctly saved and loaded in the MongoDB and no errors or exceptions happened, but at the same time, if something did go wrong, I would still have the correct data stored on disk as all operations and modifications were also applied to those files.
As soon as everything is stable with the MongoDB version, removing the old system is very easy. In the end the migration was as easy and as smooth as I expected. There is still some data that is present in both systems, and some isn't stored in MongoDB yet because I used a timestamp as a key and for some reason, trying to read it again fails using the Spring MongoDB connector. In hindsight, using a timestamp as a key in a JSON structure was a bad idea regardless.
The main struggle I had, and honestly still have, is that I am not familiar with the query language for MongoDB. It is nothing like SQL, which makes sense, but it can be pretty difficult to figure out how to write it and make it work. This is still something I should spend time on learning, and perhaps some better IDE to assist me with it, since the Mongo shell isn't the easiest option.
Afterwards, I did find out that PostgreSQL actually has pretty decent support of JSON as well. I aim to have some comparison between MongoDB and PostgreSQL on handling JSON data. I don't know what to expect with that, but one worry I have is that MongoDB has a maximum size on the content, which worried me from the beginning, especially if you are dealing with highly dynamic data where the structure and size is out of your control and depends more on the user. I don't know whether such a limitation would be present in PostgreSQL and whether the limitation would even be a real life problem.