Beijing to Establish a Dedicated Scientific Data Center

Deep News
Jul 13

At the Huairou Science City, the "Science Data Bank" established by the Computer Network Information Center of the Chinese Academy of Sciences is operating at full capacity, providing stable and continuous services to researchers worldwide. The Beijing Municipal Science and Technology Commission and the Administrative Committee of Zhongguancun Science Park have recently indicated that Beijing currently possesses a concentrated and advantageous pool of scientific data resources. The next step involves plans to construct a scientific data center in the city, alongside developing an intelligent data processing and dataset system.

The Five-hundred-meter Aperture Spherical radio Telescope (FAST), known as the "China Sky Eye," has already discovered over a thousand new pulsars and accumulated a significant batch of crucial observational data. Leveraging the "Science Data Bank," FAST has released datasets on repeating fast radio burst sources, which have supported global scientists in producing more than 300 research papers since 2021.

Scientific data encompasses the vast observational outputs from major scientific facilities, foundational research data used to form academic papers, and results data from research projects. It includes all types of digital information generated during scientific activities, such as original records, experimental measurements, observational results, and computational analyses, serving as core material for verifying discoveries and driving innovation. Currently, outstanding data achievements from fields including astronomy, brain science, and artificial intelligence, including those from FAST, are being released globally through the "Science Data Bank."

"Through a data community model, we have established data circulation links with research teams. Using pre-defined data interfaces, data generated by facilities and research teams can be automatically submitted to the 'Science Data Bank,'" explained Li Chengzan, the R&D lead for the Science Data Bank at the Computer Network Information Center of the Chinese Academy of Sciences. The platform provides data community services, allowing high-quality data to be easily integrated and rapidly enter global dissemination channels. "Scientists also hope to make data globally visible, which helps enhance the academic influence of their facilities and research teams," he added.

Currently, the "Science Data Bank" hosts high-quality scientific data resources uploaded and published by researchers from over 100 countries worldwide. Including the "Science Data Bank," 17 out of 20 National Science Data Centers are located in Beijing. These centers are continuously expanding and accumulating various scientific research data resources from projects, journal articles, and major scientific facilities.

This massive volume of scientific data is aiding researchers in generating new discoveries. "Relying on the scientific data aggregated by the National Microbiology Science Data Center, we have constructed a dedicated large scientific model for protein function mining. It can not only visualize the structure of proteins but also analyze subtle differences such as local charge distribution and hydrophobicity," said Zhang Shouyue, a researcher at the Institute of Microbiology of the Chinese Academy of Sciences and Chief Scientist at the National Microbiology Science Data Center. Using this model, the team recently discovered a novel enzyme with potential for industrial applications. This enzyme can synthesize a key precursor required for diterpenoid compounds in a single step, and under experimental conditions, the production of the target precursor increased by over 10% compared to the control group.

Without big data and large models, scientists would often need to conduct experimental screenings one by one from a vast pool of candidates, a process that is time-consuming and costly. "A particular strength of our model is its ability to distinguish subtle differences between 'twin' enzymes," Zhang Shouyue noted. Some enzymes have very similar overall sequences and structures, but minor changes in a few amino acids within the active site or the surrounding physicochemical environment can determine whether they produce different products. The model can prioritize candidates more likely to possess the target function, significantly narrowing the screening scope through experimental validation and thereby improving the efficiency of discovering target enzymes.

High-quality scientific data can be repeatedly verified and reused, greatly accelerating the research process. The recently issued "Beijing Implementation Plan for Accelerating AI-Enabled Scientific Research (2026-2028)" proposes consolidating the foundation of scientific data, advancing the construction of the Beijing Scientific Data Center, and establishing an intelligent data processing and dataset system. The Beijing Municipal Science and Technology Commission and the Administrative Committee of Zhongguancun Science Park stated that related work will efficiently unlock the value of data, enabling it to better empower fundamental research and AI-driven scientific innovation.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10