Chapter 34: Data Fusion Anomaly
Several of the senior students in the lab clicked on the link, eager to examine Zhou Yun’s secret manual. Even Qiu Yan, who had just slumped over, sat back up. After browsing for a while, they couldn't help but exclaim.
"I’d call this the Ultimate Guide for Graduate Student Beginners!"
"Indeed. If we’d had this back then, we would’ve saved so much time. For starters, just setting up the environment cost me a month or two. Learning GitHub took another week or two, and figuring out how to search for academic literature took even more. By the time you learn all these miscellaneous things, half a semester is gone. There wasn't any systematic tutorial either; we had to find everything bit by bit online. You guys are lucky—with this resource from Zhou Yun, you’ll save a lot of time."
"As long as it’s helpful to you, that’s all that matters. If anyone else needs it, just pass it along, but please don't let anyone charge for my work. If you find it helpful after reading, please leave a star."
"Already done. Mark my words, this thing is going to go viral sooner or later!"
"I’ll take your word for it."
After the brief commotion, the lab returned to silence. Zhou Yun stared at the experimental logs on his screen, feeling a sense of difficulty for the first time. He had finished the core code to support the model last week, set up several experiments, and let them run for six days. Today, the results were finally in.
However, the results were somewhat disappointing. Even when using the same stocks, the performance wasn't as good as the previous, stripped-down version that only accepted numerical and text data. This is one of the inherent problems in AI: the model is a complete black box. You never know exactly how your data is being transformed or transmitted within it; a single line of code can lead to all sorts of bizarre issues.
Fortunately, Zhou Yun had included extensive debugging code, as each experiment took too long to run. Even this time, he hadn't used the full dataset—just a portion of it—and a single batch still took a week, even with a cluster of 64 H100 GPUs. If he used all the data, it would take at least two weeks. But that was only for the initial training; once the model was trained, any subsequent new data wouldn't require retraining from scratch—just fine-tuning.
He was currently reviewing the output logs to pinpoint where things went wrong. To measure the model's effectiveness, he had designed several indicators across data preprocessing, data fusion, model training, and output. After careful observation, he identified the most likely culprit: an anomaly in the data fusion process.
Since the model accepted multi-modal data, there was a data fusion stage after preprocessing. According to the logs, the issue lay right there. The original fusion algorithm worked well with only two modalities, but as the number of modalities increased, previously hidden bugs began to surface. This was the primary reason the final performance was lagging. It could also be due to overfitting or data leakage—common issues—but based on the logs, that seemed less likely.
"Hmm... excessive variance in feature dimension contribution?" Zhou Yun’s finger stopped scrolling as he spotted an unusual output. In plain terms, the model lacked a sense of priority when fusing information; it treated all input modalities equally. This might work when there are only a few modalities, as there is often an implicit, manual filtering step before the data is fed in. For instance, if you’re predicting stock trends, you might trust financial indicators more than expert video analysis, so you subconsciously choose to feed the numerical indicators into the model rather than the videos. This implicitly assigns weight to the data, even if it isn't explicitly coded. But manual filtering has its flaws; in a field as counterintuitive as finance, experience alone often leads to wrong judgments.
"In other words, there’s a missing 'intelligent selection' step during data fusion to let the model know which data is important and which isn't."
"Data selection..." Zhou Yun tapped his fingers on the desk, pondering a solution. Simple logical judgments wouldn't work; they were too rigid and no better than human filtering. Confidence scores? Zhou Yun thought about it, but dismissed it. Confidence scores are just a measure of how certain the model is about its output—like in a classification task where the Softmax function turns outputs into probabilities. While it sounds profound, it’s actually just another rigid logical judgment.
There were many other ways to filter data, but Zhou Yun wasn't satisfied with any of them; they lacked the "intelligence" he was aiming for. Suddenly, his fingers paused. From another perspective, data selection could be viewed as a form of data distillation. Data distillation is essentially the process of purifying a dataset. As it happened, the *AgileEdge* paper Zhou Yun had published at NeurIPS included a data distillation method. Since shrinking a model meant reducing the number of parameters, the two concepts shared a similar underlying logic.
He couldn't just copy it directly, but he felt that with a few modifications, it would achieve the desired effect. After all, when he originally designed that distillation method, he had focused on the very concept of "intelligence." He pulled up the code from his previous paper, copied the encapsulated distillation method, and began modifying it to fit the current model. Since the code wasn't voluminous, he didn't use AI—the AI might not accurately grasp his intended changes, and it was better to do it himself.
By nearly six o'clock in the evening, he leaned back in his chair and stretched. Finally, it was done. Scrolling through the mouse wheel and watching the code run successfully, a wave of profound satisfaction washed over him. This was the true thrill of scientific research—the pleasure of solving a difficult problem, a feeling that nothing else could replicate.
After waiting a moment to confirm that the first epoch had started smoothly, Zhou Yun disconnected from the server. He had set up several experiments, which would likely take two weeks to complete, but he could take this time to relax. If things went well, the previous matter would also be coming to an end soon.