SMS AutomationZapiern8n

SMS message doubles in cost with one accent

When SMS includes a character outside the GSM-7 basic alphabet (an accent, emoji, or special char), the entire message switches from GSM-7 encoding (160 characters per segment) to UCS-2 encoding (70 characters per segment). A message that's 160 ASCII characters plus one accented letter becomes two billable segments, costing twice as much, because platforms like Twilio and Nexmo use GSM-7 by default for speed and cost, but flip to UCS-2 the moment a single non-ASCII character appears.

Alexey YushkinFounder, GENERAL INFORMATICS2 min read

When you send an automated SMS with a customer's name from your CRM, the message might cost twice as much as you expected. A single accent mark in a name like José or María causes the entire SMS to re-encode from GSM-7 to UCS-2, dropping the character limit per segment from 160 to 70, and turning a single message into two billable segments.

SMS has two character sets, and platforms switch between them

GSM-7 (the default) is the Basic Latin alphabet: A through Z, 0 through 9, and a short list of punctuation (space, period, comma, question mark, hyphen, parentheses, and a few others). It fits 160 characters in a single SMS segment.

UCS-2 is Unicode, the full character set. It includes accented letters (á, é, ñ, ü), emoji (😀), curly quotes ("), and symbols like € and £. It fits 70 characters per segment.

Most SMS platforms (Twilio, Nexmo, Vonage) default to GSM-7 because it's cheaper and faster. The platform scans your message for any character outside the GSM-7 set. If it finds even one non-GSM character, it switches the entire message to UCS-2.

So a message that is "Hello José, your order is ready. Click this link to confirm delivery: https://shop.example.com/confirm?id=12345" is 123 ASCII characters plus one é. That's one GSM-7 segment. But the é triggers a switch, and the same message becomes UCS-2, now 124 characters in an encoding that only fits 70 per segment. The message splits into two segments, and you're charged twice.

Per-platform behavior differs

Twilio documents this: messages default to GSM-7, switch to UCS-2 if any character falls outside GSM-7. Each segment is one SMS charge, so a 160-character ASCII message costs one unit. A 161-character message with one accented character costs two units.

Nexmo/Vonage works the same way: check the message, detect the character set, split by segment size. The per-message charge depends on segment count.

AWS SNS takes a different approach: it uses UCS-2 for all messages, so there is no switch. A 70-character message costs one unit, a 71-character message costs two. The character set is always consistent, but the cost per character is higher because you lose the 160-char GSM-7 advantage.

Google Cloud Message pricing depends on the underlying provider, so check your configuration.

The dangerous case: imported names and data

The problem surfaces when you pull names from a CRM that have not been normalized. A HubSpot lead named João, an Airtable contact with Müller, a Salesforce opportunity owner with García. Each one triggers the switch, and if your automation sends 5,000 SMS per month, the cost difference is real.

The switch also happens on emoji and special characters. A customer using a quote mark in their company name ("Tom & Jerry's Shop"), a phone number with parentheses in the source, or a date written with curly quotes (2026-10-02 instead of 2026-10-02) will trigger UCS-2. The trigger is per-message, not per-character, so one emoji costs all 160-to-70 compression for the entire SMS.

How to fix it

Option 1: Strip accents at the input boundary. When data enters your automation, normalize it to ASCII. Replace é with e, ñ with n, ü with u. This is a one-time step at the webhook entry point, before the message is composed. The risk is data loss (José becomes Jose), but for SMS address headers and order confirmations, that trade-off is usually fine.

Option 2: Pre-test the message length. Before sending, call your SMS provider's length-calculation API or a client-side library. Twilio offers a snippet, Nexmo documents the math. If the message is over 160 characters, you know it will split, and you can either shorten it or charge the customer the double cost. Logging the split count lets you audit the cost.

Option 3: Use a provider that defaults to UCS-2. AWS SNS uses UCS-2 for all messages, so you know the cost upfront. The per-character rate is higher, but the surprise is gone. Choose this if you have data you can't sanitize and cost predictability matters more than per-message price.

Option 4: Separate the name from the template. Instead of interpolating the name into the message text, send it separately or use a placeholder. "Hi [CUSTOMER_NAME], your order is ready" stays ASCII if the placeholder stays under 160 characters. This is the safest but requires discipline in the automation design.

What to log

Every SMS send should log the character count (pre-encoding) and the segment count (post-encoding) so you can audit splits later. If you see sudden jumps in message segments sent, you know a data import introduced non-ASCII characters. The per-segment rate is where the cost lives, not the per-message rate.

Start by running a sample send with one message that includes a name from your CRM. Count characters in the message, call your SMS provider's segment-calculator, and compare. If a 155-character message with an accented name shows two segments, you know the switch has triggered.

To reduce surprises, test each SMS provider's behavior with the same name: José. Send it through Twilio, Nexmo, and AWS SNS, and log the segment count. That one test is your baseline for how much the switch costs you.

For help automating SMS at scale and handling character encoding across your workflow, contact us.

Frequently Asked Questions

SOURCES & CITATIONS

  1. SMS Message Length and Character Encoding — Twiliohttps://www.twilio.com/docs/glossary/what-sms-character-limit
  2. SMS Character Encoding and Message Length — Vonage (Nexmo)https://www.vonage.com/resources/sms-character-encoding/
  3. ITU-T Recommendation T.50 - International Reference Alphabet — ITUhttps://www.itu.int/rec/T-REC-T.50/

About Alexey Yushkin

Alexey is the founder of GENERAL INFORMATICS LLC. He designs and ships AI and automation systems for businesses and operators across the US.

Connect on LinkedIn

Related reading

Want this kind of system in your business?

We build practical AI and automation systems for operators. Send us your current workflow and we will show you what to automate first.