[1] VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]//Proceedings of the 30th International Conference on Neural Information Processing Systems. Red Hook, USA: Curran Associates Inc. , 2017: 6000-6010.
[2] DEVLIN J, CHANG M W, LEE K, et al. BERT: pre-training of deep bidirectional Transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Stroudsburg, USA: Association for Computational Linguistics, 2019: 4171-4186.
[3] KOJIMA T, GU S S, REID M, et al. Large language models are zero-shot reasoners[C]//Proceedings of the 35th International Conference on Neural Information Processing Systems. Red Hook, USA: Curran Associates Inc. , 2022: 22199-22213.
[4] VATSAL S, DUBEY H. A survey of prompt engineering methods in large language models for different NLP tasks[EB/OL]. [2026-01-05]. https://arxiv.org/abs/2407. 12994.
[5] 王东清, 芦飞, 张炳会, 等. 大语言模型中提示词工程综述[J]. 计算机系统应用, 2025, 34(01): 1-10.
WANG D Q, LU F, ZHANG B H, et al. Survey on prompt engineering in large language model[J]. Computer Systems & Applications, 2025, 34(01): 1-10.
[6] SUN X, LI X, LI J, et al. Text classification via large language models[C]//Findings of the Association for Computational Linguistics: EMNLP 2023. Stroudsburg, USA: Association for Computational Linguistics, 2023: 8990-9005.
[7] LIU C, ZHANG H, ZHAO K, et al. LLMEmbed: rethinking lightweight LLM’s genuine function in text classification[C]//Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, USA: Association for Computational Linguistics, 2024: 7994-8004.
[8] JIANG Z, TANG R, XIN J, et al. Inserting information bottlenecks for attribution in Transformers[C]//Findings of the Association for Computational Linguistics: EMNLP 2020. Stroudsburg, USA: Association for Computational Linguistics, 2020: 3850-3857.
[9] GURNEE W, TEGMARK M. Language models represent space and time[EB/OL]. [2026-01-23]. https://arxiv.org/ abs/2310.02207.
[10] SKEAN O, AREFIN M R, ZHAO D, et al. Layer by layer: uncovering hidden representations in language models [EB/OL]. [2025-12-28]. https://arxiv.org/abs/2502.02013.
[11] 刘晓明, 李丞正旭, 吴少聪, 等. 文本分类算法及其应用场景研究综述[J]. 计算机学报, 2024, 47(06): 1244-1287.
LIU X M, LI C Z X, WU S C, et al. A survey of text classification algorithms application scenarios[J]. Chinese Journal of Computers, 2024, 47(06): 1244-1287.
[12] Kim Y. Convolutional neural networks for sentence classification[C]//Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, USA: Association for Computational Linguistics, 2014: 1746-1751.
[13] SONG X, PETRAK J, ROBERTS A. A deep neural network sentence level classification method with context information[C]//Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, USA: Association for Computational Linguistics, 2018: 900-904.
[14] YAO L, MAO C, LUO Y. Graph convolutional networks for text classification[C]//Proceedings of the 33rd AAAI Conference on Artificial Intelligence. Palo Alto, USA: AAAI Press, 2019: 7370-7377.
[15] KE P, JI H, LIU S, et al. SentiLARE: sentiment-aware language representation learning with linguistic knowledge[C]//Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, USA: Association for Computational Linguistics, 2020: 6975-6988.
[16] GUNEL B, DU J, CONNEAU A, et al. Supervised contrastive learning for pre-trained language model fine-tuning[EB/OL]. [2026-01-05]. https://arxiv.org/abs/ 2011.01403.
[17] BROWN T, MANN B, RYDER N, et al. Language models are few-shot learners[C]//Proceedings of the 33rd International Conference on Neural Information Processing Systems. Red Hook, USA: Curran Associates Inc. , 2020: 1877-1901
[18] LIU P, YUAN W, FU J, et al. Pre-train, prompt, and predict: a systematic survey of prompting methods in natural language processing [J]. ACM Computing Surveys, 2023, 55(9): 1-35.
[19] 顾勋勋, 刘建平, 邢嘉璐, 等. 文本分类中Prompt Learning方法研究综述[J]. 计算机工程与应用, 2024, 60(11): 50-61.
GU X X, LIU J P, XING J L, et al. Text classification: comprehensive review of prompt learning methods[J]. Computer Engineering and Applications, 2024, 60(11): 50-61.
[20] SCHICK T, SCHÜTZE H. Exploiting cloze-questions for few-shot text classification and natural language inference[C]//Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. Stroudsburg, USA: Association for Computational Linguistics, 2021: 255-269.
[21] WU H, ZHANG Y, HAN Z, et al. CoT-driven framework for short text classification: enhancing and transferring capabilities from large to smaller model[J]. Knowledge-Based Systems, 2025, 311: 113057.
[22] ZENG M, LIU C, ZHANG S, et al. Data quality enhancement on the basis of diversity with large language models for text classification: uncovered, difficult, and noisy[C]//Proceedings of the 31st International Conference on Computational Linguistics. Stroudsburg, USA: Association for Computational Linguistics, 2025: 4704-4714.
[23] LAN Z, CHEN M, GOODMAN S, et al. ALBERT: a lite BERT for self-supervised learning of language representations[EB/OL]. [2026-01-05]. https://arxiv.org/ abs/1909.11942.
[24] CLARK K, LUONG M T, LE Q, et al. Pre-training Transformers as energy-based cloze models[C]// Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, USA: Association for Computational Linguistics, 2020: 285-294.
[25] HE P, LIU X, GAO J, et al. DeBERTa: decoding- enhanced BERT with disentangled attention[EB/OL]. [2026-01-05]. https://arxiv.org/abs/2006.03654.
[26] RADFORD A, NARASIMHAN K, SALIMANS T, et al. Improving language understanding by generative pre-training[EB/OL]. [2026-01-05]. https://openai.com/ index/language-unsupervised/.
[27] LLAMA TEAM. The Llama 3 herd of models[EB/OL]. [2026-01-05]. https://arxiv.org/abs/2407.21783.
[28] QWEN TEAM. Qwen2.5 technical report[EB/OL]. [2026-01-05]. https://arxiv.org/abs/2412.15115.
[29] GEMINI TEAM. Gemini: a family of highly capable multimodal models[EB/OL]. [2026-01-05]. https:// arxiv.org/abs/ 2312.11805.
[30] RAFFEL C, SHAZEER N, ROBERTS A, et al. Exploring the limits of transfer learning with a unified text-to-text Transformer[J]. Journal of Machine Learning Research, 2020, 21(140): 1-67.
[31] LEWIS M, LIU Y, GOYAL N, et al. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension[C]// Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, USA: Association for Computational Linguistics, 2020: 7871-7880.
[32] TAY Y, DEHGHANI M, TRAN V Q, et al. UL2: unifying language learning paradigms[EB/OL]. [2026-01-05]. https://arxiv.org/abs/2205.05131.
[33] ZHANG B, CHANG K, LI C. Simple techniques for enhancing sentence embeddings in generative language models[C]//International Conference on Intelligent Computing. Singapore: Springer Nature Singapore, 2024: 52-64.
[34] HOSSEINI E, FEDORENKO E. Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language[C]//Proceedings of the 36th International Conference on Neural Information Processing Systems. Red Hook, USA: Curran Associates Inc. , 2023: 43918-43930.
[35] SOCHER R, PERELYGIN A, WU J, et al. Recursive deep models for semantic compositionality over a sentiment treebank[C]//Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, USA: Association for Computational Linguistics, 2013: 1631-1642.
[36] ZHOU Y, LI J, CHI J, et al. Set-CNN: a text convolutional neural network based on semantic extension for short text classification[J]. Knowledge-Based Systems, 2022, 257: 109948.
[37] LI K, KANG C. Deep feature extraction with tri-channel textual feature map for text classification[J]. Pattern Recognition Letters, 2024, 178: 49-54.
[38] WANG Y, WANG C, ZHAN J, et al. Text FCG: fusing contextual information via graph learning for text classification[J]. Expert Systems with Applications, 2023, 219: 119658.
[39] SUN G, CHENG Y, KONG K, et al. Text classification based on label data augmentation and graph neural network[J]. IEEE Transactions on Industrial Informatics, 2025, 21(5): 3966-3975.
[40] ZANG W, MA W, CHEN Y, et al. TextCG: a text classification framework based on information fusion via cross-graph attention network and gated recurrent unit[J]. Information Sciences, 2025: 122413.
[41] JIAO D, LIU Y, TANG Z, et al. SPIN: sparsifying and integrating internal neurons in large language models for text classification[C]//Findings of the Association for Computational Linguistics: ACL 2024. Stroudsburg, USA: Association for Computational Linguistics, 2024: 4666-4682.
[42] JIN M, YU Q, HUANG J, et al. Exploring concept depth: how large language models acquire knowledge and concept at different layers?[C]//Proceedings of the 31st International Conference on Computational Linguistics. Stroudsburg, USA: Association for Computational Linguistics, 2025: 558-573.
[43] SKEAN O, AREFIN M R, SHWARTZ-ZIV R. Does representation matter? exploring intermediate layers in large language models[EB/OL]. [2026-01-23]. https:// arxiv.org/abs/2412.09563.
|