Tech News
← Home  ·  All topics

Tokenization

5 GoKawiil briefs on this topic

Researchers develop method to convert existing LLMs to byte-level processing

A research team has proposed a technique for retrofitting pretrained subword-based language models so they operate directly on raw UTF-8 bytes rather than tokenized text. The approach builds on existing subword-trained models instead of training byte-level models from scratch, addressing a gap between prior academic claims and the lack of real-world adoption of byte-level LLMs.

Researchers retrofit token-based LLMs to read text at the byte level

A study published in Nature by Minixhofer and colleagues introduces 'byteification,' a method for converting existing token-based large language models so they can process text as individual bytes rather than whole word fragments. The authors report that byteified models retain the ability to identify individual characters within words while still matching the performance of standard token-based models.

New tool compiles fonts that render each LLM token as equal width

A web-based compiler combines any base font (like Inter or Roboto) with a chosen tokenizer (such as DeepSeek, GPT tokenizers, or Llama 3) to produce a 'token-space font' where every token occupies the same visual width. Users can preview text in the browser, upload custom fonts or tokenizers via Hugging Face links, and export compiled fonts, with an optional 'lite mode' for smaller file sizes.

Why Apple Pay and Google Wallet beat plastic cards on security

Digital wallets like Apple Pay require device authentication—Face ID, Touch ID, or a passcode—before any purchase can be confirmed, unlike many physical cards that skip verification for small transactions. They also use tokenization, replacing your actual card number with a unique, single-use code so merchants never see real card details, and they sidestep skimming devices since no card is physically inserted.

New open-source library gpu-lexer brings AI-based syntax highlighting to WebGPU

A developer released gpu-lexer, a 27.5KB JavaScript package that uses a small WebGPU-run neural model to perform syntax highlighting without relying on language-specific grammars. It tokenizes code into words, whitespace and symbols, then predicts each token's category—such as keyword, string or function—based on surrounding context, aiming to work even on languages it has never seen before.