スキル一覧に戻る
cuba6112

torchtext

by cuba6112

0🍴 0📅 2025年12月26日
GitHubで見るManusで実行

SKILL.md


name: torchtext description: Natural Language Processing utilities for PyTorch (Legacy). Includes tokenizers, vocabulary building, and DataPipe-based dataset handling for text processing pipelines. (torchtext, tokenizer, vocab, datapipe, regextokenizer, nlp-pipeline)

Overview

TorchText is a legacy library for NLP in PyTorch. While it is in a maintenance phase, it remains a common tool for handling classic NLP datasets and building vocabularies via DataPipes.

When to Use

Use TorchText for maintaining legacy NLP projects or when utilizing its built-in DataPipe-based datasets. For new projects, transitioning to native PyTorch or other modern NLP libraries is recommended.

Decision Tree

  1. Are you starting a new NLP project?
    • CONSIDER: Using Hugging Face or native PyTorch instead of TorchText.
  2. Do you need a high-performance tokenizer for production?
    • USE: RegexTokenizer and compile it with torch.jit.script.
  3. Are you using DataPipes with multiple workers?
    • ENSURE: Use a proper worker_init_fn in the DataLoader to avoid data duplication.

Workflows

  1. Building a Text Processing Pipeline

    1. Initialize a tokenizer (e.g., BERTTokenizer).
    2. Construct a Vocab object using build_vocab_from_iterator from a dataset.
    3. Create a pipeline using transforms.Sequential containing: Tokenizer -> VocabTransform -> AddToken -> Truncate -> ToTensor.
    4. Pass raw strings through the pipeline to get padded tensors.
  2. Using Built-in NLP Datasets

    1. Import a dataset from torchtext.datasets (e.g., IMDB, AG_NEWS).
    2. Initialize the DataPipe for the desired split ('train', 'test').
    3. Setup a DataLoader with shuffle=True and a proper worker_init_fn.
    4. Iterate through the DataPipe to get (label, text) pairs.
  3. Custom Regex Tokenization

    1. Define a list of regex patterns and their replacements.
    2. Instantiate RegexTokenizer with the patterns.
    3. Optionally use torch.jit.script to compile the tokenizer for production.
    4. Apply the tokenizer to raw strings to generate tokens.

Non-Obvious Insights

  • Maintenance Status: Development of TorchText stopped as of April 2024 (v0.18), marking it as a legacy library.
  • Data Duplication Risk: DataPipe-based datasets require explicit handling in the DataLoader (via worker initialization) to ensure that multiple workers don't serve the same data shards.
  • Inference Speed: Many transforms like BERTTokenizer are reimplemented in TorchScript, allowing for high-performance inference without a full Python runtime.

Evidence

Scripts

  • scripts/torchtext_tool.py: Example of building a vocabulary and tokenizer pipeline.
  • scripts/torchtext_tool.js: Node.js interface for invoking TorchText pipelines.

Dependencies

  • torchtext
  • torch
  • torchdata

References

スコア

総合スコア

50/100

リポジトリの品質指標に基づく評価

SKILL.md

SKILL.mdファイルが含まれている

+20
LICENSE

ライセンスが設定されている

0/10
説明文

100文字以上の説明がある

0/10
人気

GitHub Stars 100以上

0/15
最近の活動

3ヶ月以内に更新がある

0/10
フォーク

10回以上フォークされている

0/5
Issue管理

オープンIssueが50未満

+5
言語

プログラミング言語が設定されている

+5
タグ

1つ以上のタグが設定されている

0/5

レビュー

💬

レビュー機能は近日公開予定です