My friends and I play a word game called Undercover. If you haven't played it: most of the table gets the same secret word, say coffee. A few players get a word that's almost the same, tea, and nobody tells them they're the odd ones out: they're the Undercovers. And one player, Mr. White, gets no word at all. Everyone describes their word in turn without saying it, then the table votes someone out. Too vague and you look suspicious. Too exact and Mr. White learns the word.

We used another app for it. What I wanted was a better-looking game, with no ads, and with word pairs that are generated rather than picked from a fixed list. Halfword has the first two. The generated pairs are still to come.

What's in it

Halfword is Undercover for one phone passed round the table:

  • 4 to 20 players, one phone, iPhone or Android. No account, no internet, no ads, no tracking. Names, games and scores stay on the phone.
  • The deal: each player holds to reveal their card, reads it alone and hides it before passing on. A card can be shown again, but the whole table sees that it was.
  • Hidden roles, properly. Civilians and Undercovers get the exact same card with a different word on it, so nobody is told which one they are. Only Mr. White is told, and the table learns a player's role when they're out. A caught Mr. White gets one guess at the Civilians' word, and a right guess steals the game.
  • Three ways to vote, switchable at any moment, even mid-vote: a real secret ballot as the phone goes round, a show of hands that the host taps in, or just the result.
  • Eight special roles to mix in: the Goddess of Justice, the Lovers, the Revenger, the Duelists, the Ghost, Mr. Meme, the Boomerang and the Joy Fool.
  • Table rules: play with Mr. White or without him, and decide whether the Undercovers are hunted along with him or play for the table.
  • Points for a whole evening, on a leaderboard. More on those below.
  • Six languages: English, German, Spanish, Russian, Belarusian and Ukrainian, with 150 word pairs written for each language, never translated. The app's language and the words' language are set separately, so a mixed group can read one and play the other.
  • Shape first, colour second: a circle for the Civilians, a tilted triangle for the Undercover, an empty square for Mr. White, so colour is never the only clue.
  • Stop in the middle of a game and carry on later, right where you left off.

It's out for iPhone and Android; halfword.slnt-opp.xyz has the links to both stores.

Tonight’s table: who leads, and the next game one tap away
Tonight’s table: who leads, and the next game one tap away
Chen’s card: coffee
Chen’s card: coffee
Dario’s card is the same card, with tea on it
Dario’s card is the same card, with tea on it
Greta is Mr. White, and only she is told
Greta is Mr. White, and only she is told
Three ways to vote; here a show of hands
Three ways to vote; here a show of hands
Special roles: Chen and Fatima were the Lovers, so both are out
Special roles: Chen and Fatima were the Lovers, so both are out
The table wins, and everyone still in gets the share
The table wins, and everyone still in gets the share
Points add up over the evening
Points add up over the evening

Now the fun part: the things that took the most figuring out.

Points that slide

Points add up over a whole evening, and that's what makes them tricky. I wanted a Mr. White win to feel like one, but not to turn that player into a leader nobody could catch for the rest of the night.

The scoring I ended up with has two words in it. The Civilians, the table, play for a share. The other side, the hunted (by default Mr. White and the Undercovers), play for a stake. Both move with how many table players the table has lost, counted in steps that grow with the size of the table.

Here's the core of it, condensed from the docstring of the script that picked the numbers:

step  = (n + 2) // 6       1 at 4-9 players, 2 at 10-15, 3 at 16-20
lost  = table seats out when the game ends, by any cause
k     = lost // step
share = max(1, 6 - k // 2)   table win: to each table seat still in
stake = 9 + min(2, k)        a right guess, or the hunted outlasting the table

From 4 to 9 players, that's:

Table players lost012345
Share, each table player still in665544
Stake, each paid hunted player91011111111

So a table that loses two of its own along the way wins 5 each instead of 6, and a Mr. White who lasts that long and then names the word takes 11. Nothing is ever split, and a player who's out scores 0 from the share or the stake, even when their side wins (Mr. White's guess aside: he's out when he makes it; the Duelists and the Joy Fool add their own points on top).

Five cyan circles under a burst of lightThe table wins: the share
The smug Half between a huge rose triangle and Mr. White’s squareThe hunted outlast: the stake
Mr. White’s square under an amber speech bubbleMr. White names the word: the stake
Three endings, as the results screen draws them

The numbers weren't guessed. A simulation script, standard-library Python, plays the game as an exact Markov chain over vote states. A tie happens one round in ten, the vote finds a living Mr. White more often than anyone else, an Undercover's off clue makes them about twice as suspicious as a Civilian, and a caught Mr. White guesses right a little more often every round. Every one of those is swept over a range (243 combinations), because the point is that the answer shouldn't depend on my guesses about how people vote.

Then it plays 5,000 evenings of nine games under each scheme, including the one from the app we used to play with (10 points to Mr. White if he won, 2 to everyone still alive every game, 3 or 5 to the survivors when they won), and asks the question that matters over an evening: if someone wins the first game alone for the hunted side, how often does nobody ever catch up with them all night?

At six players:

  • the old app: 18% of evenings;
  • the first plan (the share dropping by one every step, a stake from 8 to 12): 20%, worse than the old app;
  • the scheme above: 17%.

That doesn't look like a big win until you see the other side of it. In the same model, a Mr. White seat is worth almost twice a Civilian's (1.92 times the points per game), and one hunted win is worth 4.8 of a Civilian's average games. So winning as Mr. White is rewarded noticeably, and at six players it doesn't decide the night any more often than it did in the old app.

The script also walks every reachable position and looks for anyone who would gain from voting out their own teammate. Under the final numbers that's nobody, apart from the Duelists, who are built for it, and a rare corner of the biggest tables. That check killed a few nice-sounding ideas. Splitting the stake among the winners, for example, means Mr. White gains from his own Undercover going out: in 5,950 of the 6,071 positions it checked. The old "2 points to everyone alive" went too.

And one rule fell out of all this that matters elsewhere: no score reads a ballot. An earlier version of the rules gave a point for voting for the right person, which needs to know who voted for whom. Without it, the secret ballot can be properly secret: who voted for whom stays in the engine and in the game's log on the phone, which replay needs, and nothing on any screen, at any point, can show it.

It's cheap to ask again, too: the points table comes out in under a second, the whole study on smaller samples in about two minutes, and the full run in four to five.

One codebase, two languages

Halfword is written once in Swift and SwiftUI. The Android version comes from Skip, but not the way Fare Enough's will. Fare Enough uses Skip's Fuse mode, which compiles Swift natively for Android. Halfword uses Skip Lite: the Swift is translated into Kotlin source, and Android builds that. The price is that only the Swift Skip understands can be used, and some of it means something else once it's Kotlin.

So every module's tests run twice, in two columns: as Swift on the Mac, and as Kotlin on the JVM, because Skip transpiles the tests too. Here's a taste of what the Android column caught:

  • Int is 32 bits on Android. Swift's Int becomes Kotlin's Int, so Int.max is 2,147,483,647 there, and going past it doesn't crash the way it does in Swift: it quietly wraps to a negative number. An overflow bug crashes on one side and gives a wrong answer on the other. Anything that can grow past that is Int64, and intended wrapping is spelled &+.
  • Unsigned literals don't survive. With a: UInt64, a &* 3 doesn't compile in Kotlin; a &* UInt64(3) does. UInt64(bitPattern:) doesn't compile at all. That's why the game's random generator is SplitMix64 on signed Int64, with its constants written as negative bit patterns.
  • A sort must be strict. sorted { $0 <= $1 } is fine in Swift. On Android, Java's TimSort throws "Comparison method violates its general contract!". So the rule is a strict <, with a tie-break on seat number.
  • Strings aren't equal the same way. In Swift, a precomposed é equals an e followed by a combining accent; in Kotlin it doesn't, and "👍".count is 1 on iPhone and 2 on Android. That's a real problem when Mr. White types his guess. So every guess goes through a fold table, lowercase first and then every character mapped (ä to a, ß to ss, a combining accent to nothing), so the letters any word pack uses fold the same way on both platforms. "Kaese" counts for "Käse", too, and in Russian "ё" counts as "е".
  • rounded() rounds half to even on Android. (10.5).rounded() is 11 on iOS and 10 on Android. That one matters for the pixel art, below.

My favourite: #expect(a == 2 && b == 1) transpiles to expectEqual(a == 2 && b, 1). The last == gets split off the whole expression, so with the right types it compiles and checks the wrong thing. Every rule like these got written down, with how it was checked, and there are dozens.

Android took the long way round. The first version of the engine, the word packs and the saving passed every test on both sides. Then my Mac ran out of memory under parallel builds, and I paused Android while the iPhone version shipped. When Android came back, every module's tests passed in both columns. Halfword is on Google Play now. How its Android screens were checked without an emulator is a post of its own.

A game you can replay

The rules live in a module called GameCore, and it's the strictest code in the app: it imports nothing but the Swift standard library. No Foundation, no clock, no I/O.

A game is a reducer. Every input is an action, and apply is the only way anything changes:

var game = try Game(seats: seats, config: config, pair: pair, rng: rng)
_ = try game.apply(.startGame)        // deals the cards
// … the reveals, the clues, the vote is called …
let events = try game.apply(.castBallot(voter: 2, target: 5))
// Screens draw game.snapshot: no words, no living player's role,
// never a secret ballot's target.

A refused action throws and changes nothing, because it runs on a copy. The snapshot the screens draw from has no secrets in it until the game is over, so a screen can't leak a role even by mistake. One test swaps two secret ballots and checks that nothing visible changes.

Every random thing (who's the Undercover, who gets which special role, who speaks first, who has to mime this round) is drawn from one seeded generator, in an order that is written down. It's SplitMix64, on Int64 because of the Kotlin rules above:

public mutating func next() -> Int64 {
    state = state &+ SeededRNG.gamma
    var z = state
    z = (z ^ SeededRNG.shiftRightLogical(z, 30)) &* SeededRNG.mix1
    z = (z ^ SeededRNG.shiftRightLogical(z, 27)) &* SeededRNG.mix2
    return z ^ SeededRNG.shiftRightLogical(z, 31)
}

Its outputs match Vigna's published reference values on both platforms. And since Swift's dictionaries iterate in a random order and Kotlin's in insertion order, nothing ever loops over a dictionary or a set to make a decision: seats are walked by seat number.

The payoff is that a saved game is just its seed, its setup and its list of actions. The app saves after every action. The first time it stores a game, and again when the game is over, the store replays it from exactly what it's about to write, and refuses one that doesn't come back identical. So when someone takes a call, or the phone dies halfway through a secret ballot, the game reopens exactly where it was.

The tests lean on that hard. A thousand random games are each played twice and replayed from their saved actions, and all three must agree on every event. A hundred of them are folded into one pinned number that iOS and Android both have to produce. And one test plays a whole evening of four games through the app's model, the way the screens do, killing and relaunching the app six times along the way: in the middle of a tie's revote, halfway through a secret ballot, while Mr. White is about to type his guess. It has to come back exactly right every time.

Pixels without image files

Halfword uses sixteen colours, total. Every piece of art in it starts as a text file, one character per pixel. Here's the Undercover's glyph at 16 by 16 (header shortened):

name: role_undercover_16
size: 16x16
outline: full
---
frame 0
................
..........K.....
.........KRK....
........KRRK....
.......KRRRrK...
......KRRRRrK...
.....KRRRRRrK...
....KRRRRRRrK...
...KRRRRRRRrrK..
..KRRRRRRRRrrK..
.KRRRRRRRRRrrK..
..KKRRRRRRRrrK..
....KKKKRRRRrrK.
........KKKKrrK.
............KK..
................

K is the outline, R rose, r its darker shade. A small tool, pixelkit.py (standard-library Python again), checks every file against the art rules and exports the PNGs: sizes in multiples of 8, only palette colours, the mascot drawn in its four colours and animated at exactly 8 frames a second, and no role colour on the card every player sees.

The Undercover’s glyph: a tilted rose triangle
The same glyph, each pixel drawn 6 by 6

The mascot is the Half: a small sprite whose right side never quite arrived, the pixels thinning into a dither and then into nothing. It has six moods. These are the four that loop:

idle
peek
suspicious
smug
The Half: idle, peek, suspicious and smug, at 8 frames a second

The rule for showing it is strict: whole-number scales only, 2×, 3×, 4× or 6×, one scale per screen, never smoothed, never rotated. On iPhone that's easy, because a point is exactly 2 or 3 pixels. On Android it isn't. A common phone has a density of 2.625, so "2× in dp" gives some pixels that are 5 device pixels wide and some that are 6, and pixel art falls apart when its pixels aren't the same size.

Worse, SkipUI's Image can't draw pixel art on Android at all: .interpolation(.none) is a no-op there, and every bitmap gets bilinear filtering. So the app ships no sprite as an image. It ships the same rows of palette keys and draws each sprite as SwiftUI shapes: one Path of whole-pixel rectangles per colour. On Android a source pixel is then exactly n × n device pixels, with n = round(scale × density). That round is written out as "half away from zero", because, as above, Kotlin would round 10.5 down to 10.

The cards carry the rules, too. Civilians and Undercovers mustn't know which they are, so they get one card between them, with no role colour on it. Only Mr. White's card is different, an empty square, and it's as dark as the others, so its glow can't give him away across the table:

The card back: an amber speech bubble fading out to the right
The word card: a whole amber speech bubble
Mr. White’s card: an empty square
The card back, the card every word holder gets, and Mr. White’s

The C that read as a bracket

Titles and the Half's speech are set in Micro 5, an open-source pixel font. It has two flaws: its capital C is two pixels wide with square corners, so "Chen" reads as "[hen" and "Count the votes" as "[ount the votes". Its G, with one corner cut, reads as a D: "Greta" comes out as "Dreta". Making the titles twice as large made the letters bigger, not different.

So the app ships a patched Micro 5. A script holds the new drawings and rebuilds the patched font from the upstream file, the same bytes every time: it swaps those glyphs (C and G, their accented forms and the curly quotes), then rebuilds the font's tables, bounding boxes and checksums by hand, standard library only. It keeps every advance width, because the app measures names in font pixels to make sure they fit on one title line ("Marie-Christine" is exactly 52, the limit), and a wider C would have moved that line. Micro 5 has no Reserved Font Name, so under the OFL the modified copy can keep its name.

The G is fixed: "Greta" reads as Greta. The C is better, not fixed: at full size it still looks a little like a parenthesis, "(hen". A third pixel would cure it, and push "Marie-Christine" one pixel over the line.

Chen, Count the votes and Greta, each drawn in upstream Micro 5 and then in the patched font
Upstream Micro 5, then the patched one; the dashes mark the redrawn letters

The headings of this post are set in that patched font.

Drawn first

Before the app had a line of code, it had a spec, visual guidelines, its art, and a design canvas full of screens. The guidelines set the rules this post has already leaned on (sixteen colours, whole-number scales, shape before colour), plus two more: pixel art for the content and native controls for anything you tap, and every text colour passing WCAG AA against its background.

The art-direction notes came out of drawing against those rules, and they correct them with measurements. The guidelines suggested amber-dk for amber text on the light theme; on paper it measures 3.29:1 and fails body text, so the light theme's coloured text is green-dk or rose-dk, and amber is only a fill there. Paper on rose is 3.11:1, so the labels on rose buttons are dark. The device-pixel sizing described above started as one of these notes: the guidelines' own Android snippet sized sprites in dp. And once the circle, the triangle and the square belong to the roles, no icon may use them: a triangular warning sign would read as Undercover, so the alert icon is a bare "!", and the timer is an hourglass, not a clock face. Where the notes and the guidelines disagree, the notes win.

The canvas has 44 boards. Most are phone screens at 390 by 844 points, from the splash to the leaderboard, a few of them again in the light theme. The rest set out the foundations: the palette with its measured contrast, the mascot's six states, the icons, the role art, the illustrations and the app icon. Every screen board tells the same evening, with the same eight players, coffee against tea, Dario as the Undercover, Greta as Mr. White, and scores that add up from one board to the next. Each screen declares its pixel scale, and a lint checks every image on it against that scale.

Every screen that has a board was built to match it. The app's fixtures, states it opens straight from launch for tests and screenshots, replay that same evening, and a fixture carries the name of the board it was built to match, Round-Voting or Resolve-Cascade, so a screenshot can be held up against its drawing. The store screenshots, the ones in this post among them, come from those same fixtures. Some boards are behind the rules now; until they're redrawn, the screenshots are the reference.

That's it

Halfword speaks six languages, works offline, runs on iPhone and Android, and asks nobody for an account.

And if Mr. White still steals the evening at your table, tell me. I have a script for that.