尧图网站设计 尧图网站设计YAOTU DESIGN
ARTICLE DETAIL

资讯详情

深耕网站设计与一线实操的经验洞察。

GLiNER2.5 Multi 实战指南:基于 mDeBERTa 的统一 Schema 多语言信息抽取

GLiNER2.5 Multi 实战指南:基于 mDeBERTa 的统一 Schema 多语言信息抽取 人工智能NLP信息抽取【免费下载链接】gliner2.5-multi-v1项目地址https://ai.gitcode.com/hf_mirrors/fastino/gliner2.5-multi-v1点击查看免费下载本指南以fastino/gliner2.5-multi-v1模型卡为核心系统讲解如何在一次模型中同时完成命名实体识别NER、文本分类、关系抽取、Span 属性打分与结构化记录解析。读完本文你将掌握AutoExtractor的正确加载方式、边界Boundary架构的工作机制以及从独立任务解码到联合信息抽取JointIE的完整调用链能够直接用于多语言场景下的本地推理。GLiNER2.5 Multi 是 GLiNER2 系列中的多语言边界boundary检查点基于 mDeBERTa-v3-base 构建是全系列中默认的“一个模型搞定实体、分类、记录与关系”的多语言方案。它通过稀疏的 start/end 配对代替固定宽度网格来定位任意长度的 span并借助约束解码Classifier与JointIE实现跨任务标签约束与带类型端点的实体—关系图搜索。与传统的序列标注式 NER 相比它的核心创新在于所有任务共享同一套边界打分器任务差异只体现在 Schema 的定义上。GLiNER2.5 模型家族与定位GLiNER2.5 系列包含三个检查点三者共享同一套公开 API模型参数量编码器语言适用场景fastino/gliner2.5-small-v174MDeBERTa-v3-xsmall英语快速 CPU 抽取 / 分类fastino/gliner2.5-base-v1194MDeBERTa-v3-base英语默认英语多任务检查点fastino/gliner2.5-multi-v1287MmDeBERTa-v3-base多语言默认多语言多任务检查点本仓库承载的正是fastino/gliner2.5-multi-v1的模型文件包括 config.json边界提取器完整配置、model.safetensors约 594MB 权重多为 FP16、tokenizer.json 与 tokenizer_config.jsonDebertaV2 分词器含[SEP_STRUCT]、[SEP_TEXT]、[P]、[C]、[E]、[R]、[L]等任务专用特殊 token。三个检查点使用同一套公开 API因此本文所有代码对-small-v1、-base-v1同样适用。为什么选择 GLiNER2.5一个模型多种任务实体、分类、结构化记录、关系与 Span 属性可以写在同一个 Schema里一次extract调用全部返回边界架构用稀疏的 start/end 配对代替固定 span 宽度网格只要 span 落在编码窗口内任意长度都能表示约束解码Classifier用于跨任务标签约束JointIE用于带类型端点的实体—关系图搜索本地推理通过gliner2[local]在 CPU、CUDA 或 MPS 上运行无需外部 API。边界架构原理稀疏 start/end 配对GLiNER2.5 的架构字段在仓库根目录 config.json 中明确标注为architecture: boundary、architectures: [BoundaryExtractor]。与旧版 GLiNER 的span架构在稠密[L, W]宽度网格上打分不同边界架构由两个协同工作的模块组成候选搜索先用边界打分器独立选出 top-k 的 start 与 end token配置中start_top_k: 24、end_top_k: 24、starts_per_end: 12、ends_per_start: 12再做稀疏配对而非枚举所有宽度span 打分与内容编码对配对的候选 span 计算兼容性分数pair_dim: 128、multihead_pair_compat_heads: 8并编码 span 内部内容enable_span_content: true、content_dim: 64、use_inside_evidence: true供下游分类、关系、记录任务复用。边界头boundary_head配置段还包含bidirectional_proposals: true双向候选提议、enable_abstention: true与abstention_threshold: 0.5弃权机制允许模型对不存在的实体输出“无”以及overlap_policy: flat默认用加权区间调度处理重叠 span。整个提取器支持max_len: 4096的编码窗口因此单窗口内可表示任意长度的 span但注意它并不会拼接一个首尾从未共现过的 mention。安装与加载模型安装pip install gliner2[local]需要 Python 3.10 或更高版本。[local]extra 会一并安装 PyTorch从而可以直接加载 Hub 上的检查点。通过 AutoExtractor 加载加载 GLiNER2.5 必须使用AutoExtractor。GLiNER2.from_pretrained(...)是旧版的span加载器不会分发这个边界检查点from gliner2 import AutoExtractor model AutoExtractor.from_pretrained(fastino/gliner2.5-multi-v1) print(type(model).__name__) print(model.config.architecture) # BoundaryExtractor # boundary可选设备、fp16 与编译参数model AutoExtractor.from_pretrained( fastino/gliner2.5-multi-v1, map_locationcuda, # or cpu / mps quantizeTrue, # fp16 weights on GPU compileTrue, # torch.compile after the first tracing call ) print(type(model).__name__, next(model.parameters()).device) # BoundaryExtractor cuda:0从源码结构看AutoExtractor会读取检查点config.json中的architecture字段并自动分派到BoundaryExtractor同样地Classifier、JointIE等高级组件也都基于同一个边界检查点构建内部共享实体与关系候选打分结果。因此千万不要用GLiNER2/SpanExtractor加载本检查点——它们期望的是旧版 span 架构。实体抽取基础用法text Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday. result model.extract_entities( text, [company, person, product, location], include_confidenceTrue, include_spansTrue, ) print(result) # { # entities: { # company: [{text: Apple, start: 0, end: 5, confidence: 0.98}], # person: [{text: Tim Cook, start: 10, end: 18, confidence: 0.97}], # product: [{text: iPhone 15, start: 29, end: 38, confidence: 0.96}], # location: [{text: Cupertino, start: 42, end: 51, confidence: 0.95}], # } # }返回的偏移量是半开字符区间直接作用于原始字符串text[start:end] entity[text]可放心用于下游的高亮、脱敏或文本改写。为领域标签补充描述当标签具有领域特异性时用字典传入标签描述模型会把这些描述编码为查询query显著提升抽取质量result model.extract_entities( Patient received 400mg ibuprofen for severe headache at 2 PM., { medication: Names of drugs or pharmaceutical substances, dosage: Amounts such as 400mg, 2 tablets, or 5ml, symptom: Reported symptoms or conditions, time: Clock times or relative times, }, include_spansTrue, ) print(result) # { # entities: { # medication: [{text: ibuprofen, start: 23, end: 32}], # dosage: [{text: 400mg, start: 17, end: 22}], # symptom: [{text: severe headache, start: 37, end: 52}], # time: [{text: 2 PM, start: 56, end: 60}], # } # }文本分类classify_text按任务独立解码。单标签示例result model.classify_text( This laptop has amazing performance but terrible battery life!, {sentiment: [positive, negative, neutral]}, ) print(result) # {sentiment: negative}多标签示例aspect 分析通过multi_label与cls_threshold控制输出result model.classify_text( Great camera quality, decent performance, but poor battery life., { aspects: { labels: [camera, performance, battery, display, price], multi_label: True, cls_threshold: 0.4, } }, ) print(result) # {aspects: [camera, performance, battery]}约束分类跨任务标签规则当一个任务的标签在逻辑上约束另一个任务时例如“intentdelete 必然带来 effects 含 delete”classify_text的独立解码不会强制这些规则需要改用gliner2.classification.Classifierfrom gliner2.classification import ( Classifier, ClassificationSchema, ClassificationConfig, ) from gliner2.classification import constraints as C clf Classifier.from_pretrained(fastino/gliner2.5-multi-v1) schema ( ClassificationSchema() .single(intent, [read, write, delete]) .multi(effects, [read_only, create, modify, delete], min_labels1) .constrain( C.implies((intent, delete), (effects, delete)), C.excludes((intent, read), (effects, delete)), ) ) result clf.classify(Delete the temporary file from /tmp, schema) print(result.value(intent)) print(result.value(effects)) print(result.feasible) print(result.to_dict()) # delete # [delete] # True # { # intent: { # value: delete, # confidence: 0.93, # probabilities: {read: 0.02, write: 0.05, delete: 0.93}, # }, # effects: { # value: [delete], # confidence: 0.88, # probabilities: { # read_only: 0.04, create: 0.03, modify: 0.05, delete: 0.88 # }, # }, # _meta: {feasible: True, decoder: exact}, # }这里的result.to_dict()会同时返回各标签的完整概率分布与_meta元信息result.feasible表示硬约束是否被满足。预测相关的参数解码器、beam 大小等属于调用时的ClassificationConfig而不是from_pretrainedresult clf.classify( Preview the report, schema, configClassificationConfig(decoderbeam, beam_size16), ) print(result.value(intent), result.value(effects), result.feasible) # read [read_only] True关系抽取本检查点以enable_relationsTrue训练见 config.json 中boundary_head段支持独立解码text Alice works for Acme in Paris. result model.extract_relations( text, [works_for, located_in], include_spansTrue, include_confidenceTrue, ) print(result) # { # relation_extraction: { # works_for: [{ # head: {text: Alice, start: 0, end: 5, confidence: 0.91}, # tail: {text: Acme, start: 16, end: 20, confidence: 0.91}, # }], # located_in: [{ # head: {text: Acme, start: 16, end: 20, confidence: 0.87}, # tail: {text: Paris, start: 24, end: 29, confidence: 0.87}, # }], # } # }也可以通过 Schema 指定每个关系的阈值schema model.create_schema().relations( {works_for: {threshold: 0.6}, located_in: {threshold: 0.6}} ) result model.extract(text, schema, include_spansTrue) print(result) # { # relation_extraction: { # works_for: [{ # head: {text: Alice, start: 0, end: 5}, # tail: {text: Acme, start: 16, end: 20}, # }], # located_in: [{ # head: {text: Acme, start: 16, end: 20}, # tail: {text: Paris, start: 24, end: 29}, # }], # } # }需要特别提醒独立抽取不保证类型约束——works_for的 head 不一定是人、tail 不一定是组织。若需要这类硬性保证请使用下一节的JointIE。联合信息抽取JointIEJointIE对 mention 与关系候选统一打分后在带类型端点和唯一性约束的全局图上搜索一致解from gliner2.joint_ie import JointIE, JointIEConfig joint JointIE.from_pretrained(fastino/gliner2.5-multi-v1) schema ( joint.create_schema() .entities([person, organization, location]) .relation(works_for, person, organization, unique_headTrue) .relation(located_in, organization, location) .no_self_loops() ) result joint.extract( Alice works for Acme in Paris. Bob joined Acme last year., schema, configJointIEConfig(optimizerbeam, beam_size32), ) print(result.feasible) print(result.to_dict()) # True # { # entities: [ # {id: e1, type: person, text: Alice, start: 0, end: 5, confidence: 0.94}, # {id: e2, type: organization, text: Acme, start: 16, end: 20, confidence: 0.92}, # {id: e3, type: location, text: Paris, start: 24, end: 29, confidence: 0.90}, # {id: e4, type: person, text: Bob, start: 31, end: 34, confidence: 0.91}, # ], # relations: [ # {type: works_for, head: e1, tail: e2, confidence: 0.88}, # {type: works_for, head: e4, tail: e2, confidence: 0.81}, # {type: located_in, head: e2, tail: e3, confidence: 0.86}, # ], # }注意unique_headTrue使每个 head人最多拥有一个works_for指向而Acme作为 tail 可以被多人同时指向no_self_loops()禁止实体指向自身的自环关系。务必检查result.feasibleFalse表示硬约束无法被满足这与“文本中不包含事实”是两回事for rel in result.relations: head result.entity(rel.head) tail result.entity(rel.tail) print(f{head.text} -{rel.type}- {tail.text}) # Alice -works_for- Acme # Bob -works_for- Acme # Acme -located_in- ParisSpan 属性给实体打上情感标签Span 属性是span 条件化的模型先找到实体再在这些精确 span 上为属性打分。它不是额外的实体类型也不是文档级分类。以人物情感为例from gliner2 import AutoExtractor, AttributeGroup model AutoExtractor.from_pretrained(fastino/gliner2.5-multi-v1) text ( Alice was delighted with the promotion, but Bob sounded frustrated about the delay. ) schema ( model.create_schema() .entities([person]) .entity_attributes({ sentiment: AttributeGroup( [positive, negative, neutral], applies_to[person], qualify_labelsTrue, ) }) ) result model.extract( text, schema, include_spansTrue, include_confidenceTrue, ) print(result) # { # entities: { # person: [ # { # text: Alice, # start: 0, # end: 5, # confidence: 0.96, # sentiment: {label: positive, confidence: 0.89}, # }, # { # text: Bob, # start: 44, # end: 47, # confidence: 0.95, # sentiment: {label: negative, confidence: 0.84}, # }, # ] # } # }参数说明applies_to[person]让情感属性只挂在 person 上不影响其他实体类型qualify_labelsTrue让模型侧查询编码为sentiment: positive而返回给用户的仍是短标签positive。把情感限制在人物上、同时照常抽取组织schema ( model.create_schema() .entities([person, organization]) .entity_attributes({ sentiment: AttributeGroup( [positive, negative, neutral], applies_to[person], qualify_labelsTrue, ) }) ) result model.extract( Alice praised Microsoft, but Bob criticized OpenAI., schema, include_spansTrue, include_confidenceTrue, ) print(result) # { # entities: { # person: [ # { # text: Alice, # start: 0, # end: 5, # confidence: 0.96, # sentiment: {label: positive, confidence: 0.88}, # }, # { # text: Bob, # start: 29, # end: 32, # confidence: 0.95, # sentiment: {label: negative, confidence: 0.86}, # }, # ], # organization: [ # {text: Microsoft, start: 14, end: 23, confidence: 0.97}, # {text: OpenAI, start: 44, end: 50, confidence: 0.96}, # ], # } # }注意输出差异组织 span 没有sentiment字段person span 才有——这正是applies_to约束生效的体现。结构化记录保留实例身份普通字段抽取会把所有字段拍平进互不关联的列表丢失“谁买了什么”的实例对应关系。记录模式通过anchor 字段绑定实例身份本检查点以enable_recordsTrue训练。使用natural模式并指定 anchorschema ( model.create_schema() .structure(purchase, modenatural, anchorbuyer) .field(buyer, dtypestr, cardinalityrequired_one) .field(item, dtypestr, cardinalityrequired_one) ) result model.extract( Alice bought apples and Bob bought oranges., schema, ) print(result) # { # purchase: [ # {buyer: Alice, item: apples}, # {buyer: Bob, item: oranges}, # ] # }cardinality支持required_one每个实例必须恰好一个等取值anchor 字段用于把同名不同实例的字段值正确分组到对应记录里。任务组合一次 extract 完成全部任务实体、Span 属性、分类、关系与结构化记录可以写进同一个 Schema在一次extract调用中全部返回from gliner2 import AttributeGroup schema ( model.create_schema() .entities({ person: Named people, organization: Companies or teams, product: Named products or services, }) .entity_attributes({ sentiment: AttributeGroup( [positive, negative, neutral], applies_to[person], qualify_labelsTrue, ) }) .classification(topic, [technology, business, sports, politics]) .relations([works_for, announced]) .structure(announcement, modenatural, anchorproduct) .field(company, dtypestr) .field(product, dtypestr, cardinalityrequired_one) ) text Apple CEO Tim Cook unveiled the iPhone 15 Pro for $999. result model.extract(text, schema, include_spansTrue, include_confidenceTrue) print(result) # { # entities: { # person: [{ # text: Tim Cook, # start: 10, # end: 18, # confidence: 0.97, # sentiment: {label: positive, confidence: 0.82}, # }], # organization: [{text: Apple, start: 0, end: 5, confidence: 0.98}], # product: [{text: iPhone 15 Pro, start: 32, end: 45, confidence: 0.96}], # }, # topic: {label: technology, confidence: 0.94}, # relation_extraction: { # works_for: [{ # head: {text: Tim Cook, start: 10, end: 18, confidence: 0.86}, # tail: {text: Apple, start: 0, end: 5, confidence: 0.86}, # }], # announced: [{ # head: {text: Tim Cook, start: 10, end: 18, confidence: 0.84}, # tail: {text: iPhone 15 Pro, start: 32, end: 45, confidence: 0.84}, # }], # }, # announcement: [{ # company: Apple, # product: iPhone 15 Pro, # }], # }注意语义边界文档级topic与逐人sentiment是相互独立的sentiment只挂在本条中Tim Cook的 person 实体上。批量推理texts [ Google hired Jane Doe in London., Tesla launched the Model 3 in California., ] results model.batch_extract_entities( texts, [company, person, product, location], batch_size8, include_spansTrue, ) print(results) # [ # { # entities: { # company: [{text: Google, start: 0, end: 6}], # person: [{text: Jane Doe, start: 13, end: 21}], # product: [], # location: [{text: London, start: 25, end: 31}], # } # }, # { # entities: { # company: [{text: Tesla, start: 0, end: 5}], # person: [], # product: [{text: Model 3, start: 19, end: 26}], # location: [{text: California, start: 30, end: 40}], # } # }, # ]batch_extract既可以接收单个 Schema作用于所有文档也可以接收一个 Schema 列表每篇文档对应一个 Schema。长文档处理extract(...)传入max_len会截断文本。长上下文助手则通过扫描重叠的词块并在块内映射回文档偏移来解决长文问题long_text (Quarterly overview. * 40) Satya Nadella spoke in Redmond about Microsoft. result model.extract_entities_long( long_text, [person, organization, location], chunk_size384, chunk_overlap64, include_spansTrue, ) print(result) # { # entities: { # person: [{text: Satya Nadella, start: 800, end: 813}], # organization: [{text: Microsoft, start: 837, end: 846}], # location: [{text: Redmond, start: 823, end: 830}], # } # } result model.extract_long(long_text, schema, chunk_size384, chunk_overlap64) print(result[topic]) # technology同样的思路也适用于Classifier.classify_long与JointIE.extract_long。长文档处理的限制只有 start 和 end落在同一块中的 span 才会被保留只有当关系的两个端点在同一块中都被抽取时关系才会被保留边界模型可以在一个编码窗口内表示任意长的 span但它不会拼接首尾从未在同一块中共同出现的 mention。模型细节以下信息来自本仓库的 config.json 与 encoder_config/config.json编码器 mDeBERTa-v3-base 的原始配置可直接核对架构GLiNER2boundary提取器BoundaryExtractor候选搜索稀疏 start/end 配对非稠密[L, W]宽度网格Span 长度任意不超过编码窗口的长度max_len4096编码器microsoft/mdeberta-v3-base12 层、hidden size 768、DebertaV2 结构参数量287M权重体积约 594 MB大部分为 FP16语言多语言启用的任务头分类、记录enable_recordsTrue、关系enable_relationsTrue重叠默认策略flat加权区间调度可在每次调用时用overlap_policy覆盖输入 / 输出文本 → 实体、标签、Span 属性、记录与关系边边界头的关键超参数来自boundary_head段还包括candidate_budget: 192与training_candidate_budget: 192候选预算、pool_size: 192与pool_boundary_top_k: 32候选池、ends_per_start: 12与starts_per_end: 12start/end 双向配对比例、enable_abstention: true与abstention_threshold: 0.5弃权机制、record_anchor_threshold: 0.5与record_field_threshold: 0.5记录 anchor 与字段阈值、relation_argument_proposal_threshold: 0.2关系参数提议阈值等。这些参数直接决定了推理时每个 query 的候选规模与阈值行为是理解模型运行开销的关键入口。再次强调不要用GLiNER2/SpanExtractor加载本检查点这两个类面向旧版 span 架构无法正确分发与解码。引用与许可如果在研究或产品中使用了本模型请引用misc{zaratiana2025gliner2efficientmultitaskinformation, title{GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface}, author{Urchade Zaratiana and Gil Pasternak and Oliver Boyd and George Hurn-Maloney and Ash Lewis}, year{2025}, eprint{2507.18546}, archivePrefix{arXiv}, primaryClass{cs.CL}, url{https://arxiv.org/abs/2507.18546}, }本模型以 Apache License 2.0 发布。赞分享人工智能NLP信息抽取【免费下载链接】gliner2.5-multi-v1项目地址https://ai.gitcode.com/hf_mirrors/fastino/gliner2.5-multi-v1点击查看免费下载相关推荐GLiNER2.5-Multi多语言信息抽取与微调部署跨语言场景落地完整指南GLiNER2.5 Multi多语言信息抽取与微调部署跨语言场景落地完整指南 GLiNER2.5 Multi fastino/gliner2.5 multi人工智能NLP信息抽取如何从自然语言中抽取结构化JSON记录GLiNER2.5-Multi结构记录抽取指南如何从自然语言中抽取结构化JSON记录GLiNER2.5 Multi结构记录抽取指南 GLiNER2.5 Multi 是多语言信息抽取模型一条自然语言即可提人工智能NLP信息抽取GLiNER2.5-Multi JointIE联合信息抽取全解如何构建类型化实体-关系图GLiNER2.5 Multi JointIE联合信息抽取全解如何构建类型化实体 关系图 GLiNER2.5 Multi 是一个多语言统一信息抽取模型其 J人工智能NLP信息抽取创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表