Rendered at 07:39:53 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
rdedev 4 hours ago [-]
Tabular foundation models are one of those things where when you first look into it, it does not make sense as to why they would work so well but it does.
In drug property prediction domain, tabular foundation models coupled with another foundation model for molecules are pretty close to being the state of art.
Btw the article makes heavy use of AI or is written in that way A lot of unnecessary dramatic flair that gets very tiring
3abiton 2 hours ago [-]
I would love to point them at my old time series prediction problems to see their ability. I was very proud of my lightgbm back then, performing miracles after a heck of a lot fearure engineering.
NordStreamYacht 3 hours ago [-]
True.
Why "Every number is a measured one" and not "every parameter is measured?"
Stopped reading at that point.
LLMs are like autotune. Imagine the Rolling Stones auto-tuned.
icfly2 2 hours ago [-]
Nice little test, but man this AI writing is a pain. I understand that you want to churn out blog posts, but please don't write in this breathless style. Tell your LLM that it is writing a lab report.
lyelibi 2 hours ago [-]
I have never seen tabular transformer models beat xgboost/catboost in industrial context where datasets is gigantic. They most produce these results on relatively small datasets, clearly not in the tens of millions of rows.
icfly2 2 hours ago [-]
I work with what generally qualifies as big data (a large fraction of European e commerce payments). The vast majority of this data holds no new insights. So even for xgboost the training data is trimmed down. To get these models to work one can trim the data down further, so split out by some known characteristics. For titanic (obviously a way to small dataset) split by gender and/or class.
bronlund 2 hours ago [-]
This AI slop is tiring.
Edit: I got downvoted, and it was most likely by the AI slop instigator himself. Imagine being so proud of your slop, that you take it personally when someone critiques it :D
Now I am curious as to how much of his blog is exactly like this.
Edit 2: All of it, it seems. Like; that article about Intern-Decision the 26. September 2026. The author installed it, fixed three failures, tested it on 27,256 days of weather, and published the write-up that same day as it was released.
And then manage to get cranky when someone points it out :D
not_a_feature 1 hours ago [-]
If you want a better (and Biomedical) benchmark that evaluates a whole range of models: https://tabbench-bio.eu
XGBoost’s search optimized accuracy, and afterwards I also compare by area under the curve. Which means the “fourteen of fourteen on AUC” is against a boosting model that was not tuned for that metric. Tuning it for AUC would probably improve it there; I did not measure that.
So, not a fair test?
I also did not get why Xgboost had to count its training time for the inference. You only train once. I guess in some scenario, where someone says, "I need the best model now, you have five minutes on this singular dataset", but I have never been in that situation.
I would feel better if scripts were released, because I am fairly dubious. I take it as a given that a tabular model has been pre-trained on all of the public benchmark datasets, but that is what it is.
The slop was so meandering, I am not sure what is truth or not.
In drug property prediction domain, tabular foundation models coupled with another foundation model for molecules are pretty close to being the state of art.
Btw the article makes heavy use of AI or is written in that way A lot of unnecessary dramatic flair that gets very tiring
Why "Every number is a measured one" and not "every parameter is measured?"
Stopped reading at that point.
LLMs are like autotune. Imagine the Rolling Stones auto-tuned.
Edit: I got downvoted, and it was most likely by the AI slop instigator himself. Imagine being so proud of your slop, that you take it personally when someone critiques it :D
Now I am curious as to how much of his blog is exactly like this.
Edit 2: All of it, it seems. Like; that article about Intern-Decision the 26. September 2026. The author installed it, fixed three failures, tested it on 27,256 days of weather, and published the write-up that same day as it was released.
And then manage to get cranky when someone points it out :D
or a general one: https://tabarena.ai
I also did not get why Xgboost had to count its training time for the inference. You only train once. I guess in some scenario, where someone says, "I need the best model now, you have five minutes on this singular dataset", but I have never been in that situation.
I would feel better if scripts were released, because I am fairly dubious. I take it as a given that a tabular model has been pre-trained on all of the public benchmark datasets, but that is what it is.
The slop was so meandering, I am not sure what is truth or not.