Chapter 28: The Self of Yesteryear
On Friday, Zhou Yun officially signed the research project contract with Huijin. Once the preliminary research meets the expected criteria, the transition to the full-scale project will follow.
Over the past week, both parties communicated requirements and the resources Huijin would provide. Huijin’s needs were straightforward and had been largely settled; they were now simply clarified. They required a model capable of predicting a single stock with minimal resource consumption and maximum accuracy. This served as the primary benchmark for the level of support Huijin would provide Zhou Yun moving forward.
As for resources, the deep-pocketed Huijin provided a cluster of 64 H100 GPUs, along with the corresponding processors and memory—an investment valued at over ten million. Although labeled as preliminary research, the current version was functionally similar to the final version in its core architecture, with only minor concessions in data volume and model complexity. After all, if one could accurately predict a single stock, the same logic applied to countless others; the only difference would be the scale of data. Huijin was clearly aware of this challenge, which explained their generosity.
Zhou Yun estimated these resources would be more than sufficient, as the model he was developing inherently consumed less computing power than similar models on the market. Huijin also provided a monthly stipend of 20,000 for his services. According to the contract, he had one year to complete the project; failure would result in him joining Huijin as an employee. While a year might seem tight for such a major project, Zhou Yun was confident he could deliver.
In the lab, Zhou Yun did not immediately begin designing the scheme. Instead, he started by reviewing academic papers. Both in his past life and the present, his experience had been limited to "small models"—lightweight neural networks like LSTM, CNN, or FCN, which feature simple structures and low parameter counts. The project he was now undertaking, however, required a true multimodal large model. These models are generally based on the Transformer architecture. While powerful, Transformers have a drawback: their core mechanism, "Attention," has a time complexity of O(n²), requiring immense computational power. This is why large models on the market today require thousands, or even tens of thousands, of GPUs for training.
Beyond the vast differences in computational cost and resources, large and small models also differ significantly in their overall architecture. A small model might be written and executed in just a few hundred lines of code. A true large model, however, requires not only the core logic but also a vast array of supporting functional code—tens of thousands of lines are considered modest.
Lacking experience and depth in this specific field, his first step was to read literature. He intended to master the state-of-the-art techniques in multimodal large models before planning how to upgrade his existing, low-parameter model. Ideally, this step should have been taken before the contract was signed. However, because his initial model performed so exceptionally, and because he had remained so composed during negotiations, Huijin assumed he already possessed the necessary expertise.
It was of little consequence. With his current proficiency in English and his capacity for comprehension, reading a dozen papers a day was no challenge. Within a month, he would have a solid grasp of the technology in the large model field.
By the end of July, two weeks had passed since the negotiations. On Monday, Zhou Yun arrived at the lab as usual. Just as he sat down to begin his daily reading, Shen Rui approached him, clutching a laptop.
"Junior brother Zhou Yun, I need a favor," he said, smiling sheepishly.
"Go ahead."
"Well, I submitted my draft to Professor Deng, but he hasn't been satisfied after several revisions. He says there's no innovation and no improvement in model performance, and that all my effort is in vain. I've tried all his suggestions, but there's no progress. Didn't you see me get scolded at last week's group meeting?"
He felt overwhelmed. He often wondered why he had chosen to pursue this graduate degree; compared to Zhou Yun, he felt as slow as a protozoan.
"Sure, let me take a look at the paper."
It wouldn't take long, and as colleagues, they shared a good rapport—Shen Rui often treated him to coffee, tea, and meals.
"Thank you so much! I’m sorry to bother you when you’re so busy, but I’m truly at my wit's end. If this continues, I’m certain I’ll have to delay my graduation." Shen Rui opened his laptop as he spoke.
Zhou Yun took the computer and scrolled through the document. Shen Rui’s research focused on concept drift in network traffic. Simply put, concept drift means that if WeChat traffic appears as form A on a network today, a year later—due to protocol updates or software changes—it manifests as form B. This change renders existing detection models inaccurate.
Zhou Yun was familiar with this field, having dabbled in it during his previous life while working on industrial projects for his advisor. He scanned the introduction and technical background sections, immediately grasping the context. Like most academic papers, the structure was sound. However, upon reaching the methodology section, he identified the bottleneck.
To be fair, Shen Rui’s methodology was sufficient for a CCF-C or an SCI Tier 2 journal, but Professor Deng’s standards were high, and the criticism was not entirely unwarranted. As he read, Zhou Yun couldn't help but shake his head and smile.
Seeing this, Shen Rui’s heart sank. "Zhou Yun, is there a major problem with the paper?"
He trusted Zhou Yun implicitly—in his mind, Zhou Yun’s authority was on par with Professor Deng’s. If even Zhou Yun was shaking his head, did that mean his paper was beyond saving?
"It's fine, the problem isn't that big," Zhou Yun reassured him.
He was smiling because he saw his former self in Shen Rui—lacking innate talent, forced to rely on tweaking others' models and adding modules to publish papers. There was no help for it; geniuses were rare, and most graduate and doctoral students were little more than "academic tailors."
"If you just want to get the paper published, you only need to add one module. Your model’s detection accuracy for concept drift is low because it fails to accurately identify robust features. You just need to..."
Zhou Yun thoughtfully provided the relevant papers and GitHub links. If Shen Rui followed the code and integrated the module into his own model, he would have both the innovation and the performance boost required to satisfy his advisor.