See also:
~/notes/pip.md
~/notes/pyenv.md
~/notes/pynamo.md
~/notes/numpy.md
~/notes/pandas.md
range(3, -1, -1) and slice(3, -1, -1) are very different.
You can range(3, -1, -1) to count down from 3 to zero (inclusive), but if you slice s[3:-1:-1] you always get an empty string.
In slices, -1 always means the last element in the list, regardless of the step.
s = 'abcdefg'
s[3 : 2 : -1] # => 'd'
s[3 : 1 : -1] # => 'dc'
s[3 : 0 : -1] # => 'dcb'
s[3 : -1 : -1] # => ''
Remember that -1 in start or stop of s[start:stop:step] means the last element.
And remember to put the slice args in the right direction when step is -1
"abcdef"[3:1:-1] # => 'dc' (first arg is inclusive, second is exclusive)
This pattern works great in ruby to do something if the return is not nil
if val = vend_value_or_nil
# Do something with val
end
but in python, do NOT do this if the return can be 0, empty string, or an empty container:
if val := vend_value_or_none:
# Bad! value can be 0 which is falsy
the safer version is not nearly as elegant, so maybe use another pattern entirely:
if (val := vend_value_or_none) is not None:
# Do something with val
use:
val = vend_value_or_none()
if val is not None:
# Do something with val
If I find myself ever doing something awful like this:
min_val = float("inf")
min_key = -1
for k, v in my_counter.items():
if v < min_val:
min_val = v
min_key = k
use this instead:
min_key, min_val = min(my_counter.items(), key=lambda item: item[1])
Use defaultdict(int) not defaultdict(0).
If I need a number other than 0 to start with, use defaultdict(lambda: 5)
For the 0 case, prob better to use from collections import Counter
If I find myself creating a bit of state called found and then breaking from a loop and testing:
found = False
for ...:
if meets_found_condition:
found = True
break
if not found:
do_default
Use this instead:
for ...:
if meets_found_condition:
break
else:
do_default
The else runs when the loop finishes without break, including when the loop has zero iterations.
x = 1
type(x) # => int
isinstance(x, (int, float)) # => True
class C:
pass
class D(C):
pass
isinstance(D(), C) # => True
type(D()) is C # => False
from time import time
time()
See also ~/dev/snippets/python/date_time.py
~/dev/snippets/python/get_timezone.py
~/dev/snippets/python/hours_between_two_datetimes.py
~/dev/snippets/python/time_execution.py
Instead of
from collections import namedtuple
Point = namedtuple("Point", ["x", "y"])
I can use
from typing import NamedTuple
class Point(NamedTuple):
x: int
y: int
Resize = namedtuple("Resize", ["width", "height"])
Nicer than:
points = json.loads('{"points": [{"x": 0, "y": 0}]}')
is decoding into structured objects.
First option, namedtuple:
import json
from collections import namedtuple
Point = namedtuple("Point", ["x", "y"])
data = json.loads('{"points": [{"x": 0, "y": 0}]}')
points = [Point(**p) for p in data["points"]]
Second option, dataclasses. mutable by default but can use frozen=True for immutability, or use eq=False to put them in sets:
import json
from dataclasses import dataclass
@dataclass
class Point:
x: int
y: int
data = json.loads('{"points": [{"x": 0, "y": 0}]}')
points = [Point(**p) for p in data["points"]]
Third option, pydantic, closest to Swift’s Decodable:
from pydantic import BaseModel
class Point(BaseModel):
x: int
y: int
class Payload(BaseModel):
points: list[Point]
payload = Payload.model_validate_json(
'{"points": [{"x": 0, "y": 0}]}'
)
If I already have a dictionary, use Payload.model_validate(data) instead.
Use is, it’s analogous to swift’s ===
If I want newlines:
text = """First line
Second line
Third line"""
If I don’t want newlines:
text = (
"This is a long sentence "
"continued on another source line."
)
If I want newlines but no indents:
from textwrap import dedent
dedent("""\
First line
Second line
Third line
""")
x.strip()
x.lstrip()
x.rstrip()
Or eliminate internal spaces
x.replace(' ', '')
from textwrap import dedent
text = dedent("""
Hello
world
""").strip()
# 'Hello\nworld'
yes:
long_name = 1
print(f'{long_name=}')
no:
long_name = 1
print(f'long_name={long_name}')
For vanilla pdb:
(Pdb) w # Show the call stack
(Pdb) u # Select the caller’s frame
(Pdb) list # Show the code around the caller
(Pdb) p node # Inspect that caller’s variables
(Pdb) d # Move back toward the current frame
Using a dataclass makes types not hashable
@dataclass
class Node:
value: int
{Node(1)} # => runtime error
But it does gives nodes field level equality
Node(1) == Node(1) # => true
I can use
@dataclass(eq=False)
or
@dataclass(frozen=True)
or implement __hash__ or __eq__ myself.
Frozen is best for immutable values like points that should compare equal (e.g.
(1,2) == (1,2) makes sense to be true). But frozen is not best for nodes,
because Node(1) and Node(1) I likely do not want to be equal. I want identity instead (use eq=False)
Append a ? to the fn or method I’m trying to understand
open?
Append ?? to see docstring and source!
import requests
requests??
Use the help built-in (I prefer ??):
help(sorted)
Print docstring directly:
print(open.__doc__)
Use dir, the old standby:
x = “abc”
dir(x)
Before 3.9:
from functools import lru_cache
@lru_cache(maxsize=None)
def fn...
After 3.9:
from functools import cache
@cache
def fn...
If in ipython or jupyter notebook, use magic %timeit
For a single pass, use perf_counter, which is guaranteed to be monotonic:
import time
t = time.perf_counter()
# do work
print(time.perf_counter() - t)
For repeated runs and summary statistics, use timeit:
import timeit, statistics
def work():
pass
times = timeit.repeat(work, number=1000, repeat=5)
print("runs:", times)
print("min:", min(times))
print("max:", max(times))
print("mean:", statistics.mean(times))
print("stdev:", statistics.stdev(times))
Can get a ruby-ish feel with a context manager:
import time
from contextlib import contextmanager
@contextmanager
def timed_block(label: str = "timed_block"):
start = time.perf_counter()
try:
yield
finally:
end = time.perf_counter()
print(f"{label}: {(end - start) * 1000:.3f} ms")
with timed_block():
total = sum(i * i for i in range(10_000_000))
Using the decorator:
from contextlib import contextmanager
@contextmanager
def example():
print("enter")
try:
yield "some value"
finally:
print("exit")
with example() as value:
print(value)
Or manually creating a class and implementing the two required dunder methods:
class Example:
def __enter__(self):
print("enter")
return "some value"
def __exit__(self, exc_type, exc_value, traceback):
print("exit")
return False # Do not suppress exceptions raised raised inside the with block
with Example() as value:
print(value)
It is valid to return from inside a with statement.
Do not do this:
unsafe_lock = Lock()
unsafe_lock.acquire_lock()
res = n + 1
unsafe_lock.release()
return res
Just do:
with unsafe_lock:
return n + 1
from operator import itemgetter
x = ['a', 'b', 'c']
indices = [0, 2]
itemgetter(*indices)(x) # => ('a', 'c')
Or, if I can’t remember all that
x = ['a', 'b', 'c']
indices = [0, 2]
[x[i] for i in indices]
Or
x = ['a', 'b', 'c']
[y for i,y in enumerate(x) if i in [0,2]]
from enum import StrEnum
class Role(StrEnum):
SYSTEM = "system"
USER = "user"
Role.SYSTEM.value # => "system"
from enum import Enum, auto
class Role(Enum):
SYSTEM = auto()
USER = auto()
class C:
@classmethod
def y(cls):
return cls
C is C.y() # => True
import array
array.array('I', [1,2])
from time import sleep
sleep(0.1)
Basic setup to get an async context in python:
import asyncio
async def main() -> None:
print("sleeping for one second...")
await asyncio.sleep(1)
print("done!")
asyncio.run(main())
Pydantic doesn’t use positional arguments:
from pydantic import BaseModel
class User(BaseModel):
name: str
# Create a user from safe input
user = User(name="a")
# Or
user_data = {'name': 'a'}
user = User(**user_data)
# Create a user from unknown user input
User.model_validate({'name': 1})
# Or
User.model_validate_json('{"name": "lou"}')
# Get json:
user.model_dump_json()
Primary mechanism is ref counting:
In CPython, the primary algorithm for garbage collection is reference
counting. Essentially, each object keeps count of how many references point
to it. As soon as that refcount reaches zero, the object is immediately
destroyed: CPython calls the __del__ method on the object (if defined) and
then frees the memory allocated to the object. In CPython 2.0, a
generational garbage collection algorithm was added to detect groups of
objects involved in reference cycles
Fluent Python pg 219
x = (1,)
id(x)
x += (2,)
id(x) # => different ID
But be careful when a tuple holds a mutable instance:
x = ["a"]
y = (x,)
x.append("b")
y
(['a', 'b'],)
str
tuple but not list:
x = (“a”,)
y = x
x += (“b”,)
x # => (‘a’, ‘b’)
y # => (‘a’,)
y is x # => False
versus
x = [“a”]
y = x
x += [“b”]
x # => [‘a’, ‘b’]
y # => [‘a’, ‘b’]
y is x # => True
def fn():
return "hello".count("l")
Raw bytecode:
fn.__code__.co_code
Readable bytecode:
import dis
dis.dis(fn)
Or:
list(dis.get_instructions(fn))
I’ve been using flake8, but starting to use ruff instead. E.g.
ruff check
ruff format
If I use parens, it’s just grouping, not a function call:
assert(x == 1)
Is actually
assert (x == 1)
It’s as if I were writing:
if(true):
Which I would not do.
https://pythontutor.com/render.html#mode=display
Use the /
def fn(a, /):
pass
fn(a="1") # Raises TypeError
d = {'a': 'b'}
d.pop('c', None)
This does not do what I expect:
for k, v in { 'ab': 12 }:
print(k)
print(v)
# => prints 'a' followed by 'b', because I'm actually unpacking the key only
To destructure the keys and values from a tuple at each iteration, use items():
for k, v in {'ab': 12 }.items():
print(f'k: {k}, v: {v}')
I always liked NameTuple. Pydantic went with mutable by default, which makes
sense for the domain. I can use this where I want immutability:
from pydantic import BaseModel, ConfigDict
class User(BaseModel):
id: int
name: str
model_config = ConfigDict(frozen=True)
Then:
u = User(id=1, name="foo")
u.id=2
# => Raises a ValidationError
Importing typing is no longer necessary.
But, I still need mypy as the static checker.
There is nothing built into python to static check your types.
a: tuple[int, str] = (1, "hi")
b: list[str] = ["a", "b"]
c: dict[str, str] = {"a": "b"}
d: set[int] = {1, 2, 3}
from collections.abc import Callable
e: Callable[[int], int] = lambda x: x + 1
f: Callable[[int, int], int] = lambda x, y: x + y
g: Callable[[int | None], int] = lambda x: x if x else 0
For optionals, I can either use:
# 3.10+ only
h: int | None = None
i: int | None = 1
or:
from typing import Optional
j: Optional[int] = None
k: Optional[int] = 1
or:
from typing import Union
l: Union[int, None] = None
m: Union[int, None] = 1
s = set()
s.add(1)
1 in s # => True
NamedTuples are immutable:
from typing import NamedTuple
class Point(NamedTuple):
x: int
y: int
p = Point(x=1, y=0)
p.y = 1
# => Raises AttributeError
You can put these in a set:
{ Point(0,0) }
or dictionary, and keying works how I’d expect:
d = { Point(0,0): 'a'}
x = Point(0,0)
y = Point(0,1)
d[x] # => 'a'
d[y] # => KeyError
Dataclasses are mutable:
from dataclasses import dataclass
@dataclass
class Point:
x: int
y: int
p = Point(x=1, y=0)
p.y = 1
# => Does not throw
from dataclasses import dataclass
@dataclass(frozen=True)
class D:
field = 1
d = D()
d.field = 2 # => raises FrozenInstanceError, as expected
d.field # => 1
D.field = 99
d.field # => 99
Don’t do this:
@dataclass
class D:
l = []
D().l.append('a')
D().l
# => ['a']
Do this:
from dataclasses import dataclass, field
@dataclass
class D:
l: list[int] = field(default_factory=list)
D().l.append('a')
D().l
# => []
from __future__ import annotations
from dataclasses import field
@dataclass(eq=False)
class Node:
val: int
neighbors: list[Node] = field(default_factory=list)
__post_init__ is a dataclass hook, it is not part of the standard object model.
from dataclasses import dataclass, field
@dataclass
class C:
quota: int
remaining: int = field(init=False)
def __post_init__(self):
self.remaining = self.quota
def use(self, amount: int):
self.remaining -= amount
https://docs.python.org/3/library/dataclasses.html#dataclasses.field
It’s not hashable, but it does guarantee elements can’t chance:
from types import MappingProxyType
x = MappingProxyType({'a': 1})
x['a'] = 2 # => raises 'TypeError: 'mappingproxy' object does not support item assignment'
Note a mutable type still allows mutation in the collection:
x = MappingProxyType({'a': [1]})
x['a'].append(2) # Fine
x['a'] # => [1, 2]
I always assume this is valid python, it’s not: x ? x : 0
I want: x if x else 0
Foo.__mro__
pdbpp is an absolute must. I don’t use the following any more. I just stick to pdb and pdbpp with this in ~/.pdbrc.py
from pdb import DefaultConfig
class Config(DefaultConfig):
sticky_by_default = True
No longer used, because ipdb doesn’t play super nicely with pdbpp:
`ipdb` is also handy, because I can get it show additional lines around a breakpoint.
Add this to `~/.bash_profile`
# Use ipdb when python hits a `breakpoint()`
export PYTHONBREAKPOINT="ipdb.set_trace"
# Show additional context around ipdb breakpoints
# Change this by typing 'context NUM' when a breakpoint is hit. Then 'bt' to redraw the source.
export IPDB_CONTEXT_SIZE=5
With that config, I don't have to explicitly breakpoint with `import ipdb; ipdb.set_trace(context=5)`.
I can use `breakpoint()` and then either `sticky` if I want to use pdbpp or `context NUM` followed by `bt` if I just want to see more.
python -m asyncio
https://github.com/amazonlinux/amazon-linux-2023/issues/483#issuecomment-1928446605
^ Don’t do this. Do this:
sudo dnf install python3.12
Available in python 3.8+
if (n := "world"):
print(f"{n} hello")
I prefer explicit imports, but I see this often. __all__ defines what should be imported on import * from xyz
__all__ = ["MyClass"]
class MyClass():
pass
from collections import defaultdict
x = defaultdict(str)
x['y'] # => ''
Or commonly with a list
x = defaultdict(list)
x['y'].append(1)
sorted(['b', 'a']) # => ['a', 'b']
sorted('ba') #=> ['a', 'b']
sorted(('b', 'a')) # => ['a', 'b']
Only immutable objects can be hashed and added to sets / dicts.
The trick is to use a tuple instead:
s = set()
x = ['a', 'b']
s.add(x) # => raises TypeError
s.add(tuple(x))
"a b".split(" ") # => ["a", "b"]
list("ab") # => ["a", "b"]
''.join(['a', 'b'])
Old:
from typing import Callable
New:
from collections.abc import Callable
The Callable syntax is Callable[[arg1, arg2], ret].
To support arbitrary args, use Callable[..., ret].
from typing import Protocol
from typing import TypeVar
class ExampleProtocol(Protocol):
"""
A common interface for types that provide a `foo` property, which
can be helpful to satisfy mypy when duck typing / using generics.
"""
@property
def foo(self) -> float:
...
# Later
# from whatever.example_protocol import ExampleProtocol
T = TypeVar('T', bound=ExampleProtocol)
def requires_a_foo(x: T):
print(x.foo)
x = [0, 1, 2, 3]
s = slice(None, None, 2)
x[s] = ['a', 'b']
x
# => ['a', 1, 'b', 3]
pairs = [(1, "Charlie"), (2, "Alice"), (3, "Bob")]
sorted(pairs, key=lambda pair: pair[1])
# [(2, "Alice"), (3, "Bob"), (1, "Charlie")]
nums = [1, 2, 3, 4, 5]
Instead of:
list(filter(lambda n: n % 2 == 0, nums))
Use:
[n for n in nums if n % 2 == 0]
samples = [(1, ['a']), (2, ['a', 'b'])]
Instead of:
for sample in samples:
stack = sample[1]
idx = sample[0]
pass
Use:
for idx, stack in samples:
pass
It’s cool that that works.
I was writing a sample tokenizer and did this:
elif char in ['+', '*', '(', ')']:
better to write:
elif char in '+*()':
No:
x != None
Yes:
x is not None
[list(x) for x in batched('ABCDEFG', 3)]
[['A', 'B', 'C'], ['D', 'E', 'F'], ['G']]
I always think I want something called grouped or chunked, but it’s batched!
I sometimes also confuse batched with divide (up next).
from more_itertools import divide
divide(3, 'ABCDEFG')
The bare * means all parameters after it must be passed by keyword:
def grouper(iterable, n, *, incomplete='fill', fillvalue=None)
Example from itertools recipes: docs.python.org/3/library/itertools.html
it = iter(range(10))
next(it) # => 0
from itertools import islice
gen = (x * x for x in range(1, 10))
list(islice(gen, 0, 3))
# => [1, 4, 9]
list(takewhile(lambda x: x > 2, [4,3,2,1]))
zip stops at the shortest arg, but there is a flavor in itertools:
from itertools import zip_longest
list(zip_longest('abc', [1,2], fillvalue=-1))
# => [('a', 1), ('b', 2), ('c', -1)]
import itertools
first5 = itertools.islice(gen, 5)
print(list(first5))
my_bytes.decode('utf8')
Python does not automatically bind error.
Catch specific exceptions when possible:
try:
do_work()
except MyError as err:
print(err) Use except Exception: for a broad catch of ordinary errors.
Avoid bare except: because it also catches KeyboardInterrupt and SystemExit.
The idiomatic way to do something like swift’s error cases:
enum PipelineError: Error {
case invalidInput
case unsupportedFormat(String)
}
Is to subclass Exception:
class PipelineError(Exception):
pass
class InvalidInputError(PipelineError):
pass
class UnsupportedFormatError(PipelineError):
pass
then:
raise UnsupportedFormatError("Unsupported format: TIFF")
and handle with:
try:
run_pipeline()
except InvalidInputError:
print("Check the input")
except UnsupportedFormatError as err:
print(err)
except PipelineError:
print("Some other pipeline error")
View the MRO with:
UnsupportedFormatError.mro()
# Do not need to define x here
try:
raise ValueError("broken precondition")
except:
x = 1
x += 1 # => 2
Huh, same as ruby
begin
raise 'bad'
rescue
x = 1
end
x + 1 # => 2
With js I need to define x ahead of time:
let x;
try {
throw Error("Bad");
} catch {
let x = 1;
}
x + 1; // => 2
pytest -svv
In modern versions of python breakpoint() will suffice.
No need to import ‘pdb’ explicitly.
TODO: This file needs to be formatted.
Also see ~/notes/pip.md
Also see ~/notes/pyenv.md
https://www.youtube.com/watch?v=MCs5OvhV9S4
Source: https://news.ycombinator.com/item?id=36785005
Serve current directory at http://localhost:8000
python -m http.server
After prototyping in the repl, dump all code I’ve entered into the repl:
%save my_file.py
So awesome!
Examine MRO with:
class Foo
pass
class Bar:
pass
class Baz(Foo, Bar):
pass
Baz.__mro__
https://github.com/satwikkansal/wtfpython
p file
https://docs.python.org/3.6/library/pdb.html#pdbcommand-step
unt is helpful to get out of a loop
Also use up and down
ipython --TerminalInteractiveShell.autosuggestions_provider=None
Check mypy types in the repl with:
pip install mypy_ipython
ipython
%load_ext mypy_ipython
… do some prototyping
%mypy
run tests with –pdb to drop you into a repl on test failure
Use mock.mock_calls to get a list of all calls sent to a mock.
Or mock.call_args_list. I’m not sure what the difference is.
x = ‘hello’
print(f’{x = }’)
Also see alternatives ‘ice cream’ and ‘q’: https://github.com/zestyping/q
Use bt to print frames, then
f [number]
to jump to one
PYTHONPATH=“./:$PYTHONPATH” python path/to/script.py
Install https://github.com/lukejpreston/xunit-viewer
Then
pytest –junit-xml=build/unit.xml
xunit-viewer -r build/unit.xml
pip install my-package==
import this
cd ~/dev/my_app
source .env/bin/activate
pip install -e ~/dev/my_lib
:!PYTHONPATH=“app:$PYTHONPATH” python %
.py files can be imported as a module, but only if they do not contain hyphens in the name
bad: from bad-thing import foo
good: from good_thing import foo
See ~/dev/snippets/python/example_package/README.md and the associated
package layout
sdist is a source distribution, build with:
python setup.py sdist
Build wheel distribution with:
python setup.py sdist bdist_wheel
Note that .whl files can be unzipped with unzip
The distributions are saved in the relative dist/ directory. Untar the source
distribution or unzip the .whl distribution to debug.
To include a non-python file in the wheel, the setting include_package_data
must be True in setup.py AND an include statement must be present in
MANIFEST.in
Wheel is the more modern form of egg: https://packaging.python.org/discussions/wheel-vs-egg/
> Wheel is currently considered the standard for built and binary packaging for Python.
Good slides here: https://blog.ionelmc.ro/presentations/packaging/#slide:1
Pickle is fully specified to reconstruct types
Insert breakpoint in problematic process, pickle.dump to file
pickle.load in script with debugging aids
http://calpaterson.com/mypy-hints.html
import sys
sys.getsizeof(full_dataset)
See ~/notes/random_notes_worth_keeping.txt search ‘memory usage’
Seeing: ModuleNotFoundError: No module named ‘
Fix: Add an init.py file into
pip install git+https://github.com/
http://diveinto.org/python3/porting-code-to-python-3-with-2to3.html#next
Most articles I’ve found on python mocking have been junk. Collection of easy to follow tips:
https://wesmckinney.com/blog/spying-with-python-mocks/
https://stackoverflow.com/a/25424012/143447
deactivate function$ python -v
> import
Install pdb++ https://pypi.org/project/pdbpp/
pip install pdbpp
Then at the debugger prompt, type ‘sticky’ for a much better debugging experience
type ‘interact’ at the pdb prompt
$ /usr/local/Cellar/python@2/2.7.15_2/bin/pip2.7 install virtualenv
$ python2.7 /usr/local/lib/python2.7/site-packages/virtualenv.py .env_2_7
$ source .env_2_7/bin/activate
sudo /usr/bin/easy_install-2.7 pip
/usr/bin/python2.7 -m pip install virtualenv
cd
/usr/bin/python2.7 -m virtualenv .env
source .env/bin/activate
pip install -r requirements.txt
list .
It’s tricky to add code to a previous code block after up-arrowing to it. To insert a newline, use:
ctlr+o+n or ctrl+q+j
%load filename.py
Use pp to pretty print objects. Not built in anymore?
from pprint import pprint as pp
pytest -s to show print statements when you run tests
very confusing that the default setting eats print statements
pytest -s -vv to show verbose difference
pytest -x
pytest -k ‘test_the_thing’
pytest path/to/file.py
%load_ext autoreload
%autoreload 2
Or add this to ~/.ipython/profile_default/ipython_config.py to automatically
reload on every session:
c.InteractiveShellApp.exec_lines = [‘%load_ext autoreload’, ‘%autoreload 2’]
%prun some_function
$ virtualenv .env
python -m venv .env
(add env to gitignore)
source .env/bin/activate
[optional]
see ~/notes/jupyter.txt
deactivate
pip freeze > requirements.txt
Add .env to .gitignoremy_environment
python -m venv .env
source .env/bin/activate
pip install -r requirements.txt
https://access.redhat.com/blogs/766093/posts/2592591
ipython –matplotlib
See ~/dev/snippets/python/matplotlib_experiment.py for an example
Within the plot window, use control to pan (drag along the x or y axis), or use
the zoom rect feature on the toolbar
Say you have a variable ‘n’. Use exclamation point in front:
ipdb> !n
Find all methods that have pid in it:
[x for x in dir(os) if re.search(‘pid’, x, re.I)] # => re.I for ignore case
help(s.listen)
or
print(s.listen.__doc__)
from future import print_function
Getting list of methods:
> dir(StringIO.StringIO)
Empty class implementation:
class Foo:
pass
Debugging
Get class name:
ipdb> x.__class__
or
ipdb> x.__class__.__name__
Packages
$ pip install numpy
$ pip install matplotlib
$ pip install ipdb # => debugger
Autocompletion:
$ pip install jedi
~/.vim/bundle $ git clone --recursive https://github.com/davidhalter/jedi-vim.git
Interactive python:
$ python -i
Better:
$ ipython
System python installs to
pip installs things to: /usr/local/lib/python2.7/site-packages
Pretty printing:
from pprint import pprint as pp
Longer pretty printing:
import pprint
params = {
'latitude': 37.775818,
'longitude': -122.418028,
'server_token': 'snip'
}
pp = pprint.PrettyPrinter(indent=4, width=1).pprint
pp(params)
Pretty printing something that resembles a dictionary
pp(dict(response.headers)) # => note the width of 1 above is important
Patches
### python debugger, ipdb, list, patch
In /home/lou/dragon/lib/python2.7/site-packages/IPython/core/debugger.py,
Changed context from 3 to 10:
def print_stack_entry(self,frame_lineno,prompt_prefix='\n-> ',
context = 10):
Configuring iPython
Create default profile:
$ ipython profile create
$ ipython locate profile
iPython tricks
Run a script:
> %run workbench.py
Insert an enter (newline) in repl:
ctrl+v + ctlr+j
Also try %edit
Exec from command line
python -c "print('sup')"
Sys
import sys
sys.path
Entry detection
if __name__ == '__main__':
print("yes, I am main")
slots = [0] * 26
c = 'x'
slot = ord(c) - ord('a')
slots[slot] += 1
chr(97) # => 'a'
The arguments are:
The type vended
The type pushed into the generator by send, if used
The type returned as the value property on StopIterator
from typing import Generator
def fn() -> Generator[int, None, None]:
return (i for i in range(10))
def fn2() -> Generator[int, None, None]:
yield 1
def fn3() -> Generator[int, None, str]:
yield 1
return “abc” # The only way to get this is on the value property of StopIteration
def fn4() -> Generator[str, str, None]:
name: str = yield “I am now primed”
yield f”{name} world”
from typing import Generator
def fn() -> Generator[str, str, None]:
name: str = yield "I am now primed"
yield f"{name} world"
g = fn()
# Prime the generator. The argument needs to be None, otherwise this will raise:
# TypeError: can't send non-None value to a just-started generator"
print(g.send(None))
# Now I can get the coroutine to spit out "hello world"
print(g.send("hello"))
import asyncio
async def fn(name: str) -> str:
return f"{name} world"
async def main():
result = await fn("hello")
print(result)
asyncio.run(main())
Bad:
class C:
def __init__(self):
_x = "x"
Good:
class C:
def __init__(self):
self._x = "x"
sys.getswitchinterval()
Change it with:
sys.setswitchinterval(1E-2)
Not sure how many python devs use this, but a translation of obj-c/swift delegate pattern could be:
class C:
def __init__(self, delegate):
self.delegate = delegate
def dowork(self):
self.delegate.did_finish()
class D:
def did_finish(self):
print("work is done")
c = C(D())
c.dowork() # => "work is done"
Or, with typing:
from typing import Protocol
class P(Protocol):
def did_finish(self) -> None:
...
class C:
def __init__(self, delegate: P):
self.delegate = delegate
def dowork(self) -> None:
self.delegate.did_finish()
class D(P):
def did_finish(self) -> None:
print("work is done")
c = C(D())
c.dowork()
Using callbacks with pluggable closures
from collections.abc import Callable
class C:
def __init__(self, did_finish: Callable[[], None]):
self.did_finish = did_finish
def dowork(self) -> None:
self.did_finish()
def done() -> None:
print("work is done")
c = C(done)
c.dowork()
Or
c = C(lambda: print("work is done"))
c.dowork()
Or
class X:
def __call__(self) -> None:
print("work is done")
c = C(X())
c.dowork()
from itertools import product
list(product(('a', 'b'), (1, 2)))
# => [('a', 1), ('a', 2), ('b', 1), ('b', 2)]
Or
list(product(('a', 'b'), repeat=2))
[('a', 'a'), ('a', 'b'), ('b', 'a'), ('b', 'b')]
from random import randint
randint(0, 10) # => element in [0, 10]
Or
from random import randrange
randrange(0, 10) # => element in (0, 10]
Or
from random import choice
choice(range(10))
Say I want a list of random numbers in [1, 5]. I naturally think:
map(lambda _: randint(1,5), range(10))
but it’s more idiomatic to:
[randint(1,5) for _ in range(10)]
Or, better:
from random import choices
choices(range(1, 6), k=10)
Or, by far the fastest for large N:
import numpy as np
np.random.randint(0, 11, size=100000)
Example from: https://docs.python.org/3/tutorial/classes.html
class Dog:
tricks = [] # mistaken use of a class variable
def __init__(self, name):
self.name = name
def add_trick(self, trick):
self.tricks.append(trick)
d = Dog('Fido')
e = Dog('Buddy')
d.add_trick('roll over')
e.add_trick('play dead')
d.tricks
['roll over', 'play dead']
[1,2,3].pop() # => 3 in O(1)
[1,2,3].pop(0) # => 1 in O(n)
Use deque instead (double ended queue):
from collections import deque
d = deque([1,2,3])
d.pop() # => 3 in O(1)
d.popleft() # => 1 in O(1)
deque also has automatic eviction:
d = deque([1,2,3], maxlen=3)
d.append(4)
d
# => [2,3,4]
Also remember for a known-good tree, the visited set is not necessary:
visited = {start}
queue = deque([start])
while queue:
node = queue.popleft()
for neighbor in graph[node]:
if neighbor not in visited:
visited.add(neighbor)
queue.append(neighbor)
x = [1,2].append(3)
x # => none
from bisect import bisect, bisect_left
x = [0,1,1,2]
bisect(x, 1) # => 3
bisect_left(x, 1) # => 1
from bisect import insort
insert(x, 1.5)
x # => [0,1,1,1.5,2]
bin(255)
# => '0b11111111'
from typing import TypeVar
from typing import Generic
T = TypeVar('T')
class Node(Generic[T]):
def __init__(self, value: T):
self.value = value
node1: Node[int] = Node('1') # This will throw a mypy error
node2: Node[int] = Node(1) # mypy happy
I don’t need to use Generic for this. It’s just:
from typing import TypeVar
T = TypeVar("T")
def rotate(list: list[T]) -> list[T]:
...
"hello".encode('utf-8')
bin.decode('utf-8')
And ‘utf-8’ is the default, so I can actually use:
"hello".encode()
bin.decode()
In this example I’m using rb and then decoding as utf-8.
It’s not necessary to do this, but I want to remember how to read as binary:
with open('read_in_chunks.txt', 'rb') as f:
while binary := f.read(2):
print(binary.decode('utf-8'))
from io import BytesIO
with BytesIO(b'hello world') as stream:
print(stream.read(3))
Note that just like an ascending range, the range is not inclusive of the second argument:
list(range(3, 0, -1)) # => [3, 2, 1]
!!foobool(foo) insteada = f'{15:x}' # => 'f'
b = format(15, 'x') # => 'f'
c = '{:x}'.format(15) # => 'f'
d = '{num:x}'.format(num=15) # => 'f'
__slots__More efficient, sure. But what I really like is this:
...with the default __dict__, a misspelled variable name results in the
creation of a new variable, but with __slots__ it raises in an
AttributeError.
Source: https://wiki.python.org/moin/UsingSlots
For example:
class C:
def __init__(self):
self._x = None
def run(self):
self.x = 1
c = C()
c.run() # => Nothing raised
c.__dict__ # => contains `_x` and `x`
versus:
class C:
__slots__ = ('_x',)
def __init__(self):
self._x = None
def run(self):
self.x = 1
c = C()
c.run() # => Raises "AttributeError: 'C' object has no attribute 'x'"
The setter syntax is a little funky, and relies on the existance of the getter.
class C:
@property
def p(self):
return self._p
@p.setter
def p(self, value):
self._p = value
frozenset is helpful to make things hashable, but you can’t use it blindly.
Say I was trying to use [‘a’, ‘b’] as a key to a dict. I can’t do this:
{['a', 'b']: 'x'} # => Raises "TypeError: unhashable type: 'list'"
but I could do this:
fs1 = frozenset(['a', 'b'])
d = {fs1: 'x'}
but beware (note the ordering of the array):
fs2 = frozenset(['b', 'a'])
d[fs2] # => 'x'
Use a tuple if order and duplicates matter as the key. Tuples are hashable (unlike swift!)
Wrong:
set(frozenset()) # => Not what I expect, but it makes sense. `frozenset()` is an iterable that the `set(...)` initializer iterates over to set the initial contents of the set.
Right:
set([frozenset()])
There is nothing like ruby’s built in [1, [2,3]].flatten # => [1,2,3].
For a list of lists, from_iterable works well:
from itertools chain
list(chain.from_iterable([[1], [2, 3]])) # => [1,2,3]
but for [1, [2, 3] example, I’d have to use type inspection:
from itertools import chain
l = [1,[2,3]]
chain.from_iterable(x if isinstance(x, list) else [x] for x in l)
for well structured, can also just use:
inp = [[1], [2, 3]]
[x for item in inp for x in item]
remember the trick to these nested list comprehensions is that the outer loop is the first for, e.g. for item in inp
x = re.match(r'a(b)', 'ab')
x.group(0)
# => 'ab'
x.group(1)
# => 'b'
from collections import Counter
Counter('xyzz')['z'] #=> 2
+ overloadc = Counter()
c['a'] += 1
d = Counter()
d['a'] += 1
d['b'] = 99
c + d
# => Counter({'b': 99, 'a': 2})
Two facts that make all this simpler:
r - c = d for any diagonal d that I’m iterating
There are m + n - 1 diagonals
matrix = [
[1, 3, 5, 7],
[2, 4, 6, 8],
[9, 11, 13, 15]
]
m = len(matrix)
n = len(matrix[0])
print("Iterate all diagonals:")
for d in range(-n + 1, m):
for r in range(m):
c = r - d
if 0 <= c < n:
print(f'visiting {matrix[r][c]=}')
print("Iterate diagonals that start at the top row:")
for d in range(-1, -n, -1):
for c in range(n):
r = c + d
if 0 <= r < m:
print(f'visiting {matrix[r][c]=}') m = 5
n = 3
nums = list(range(1, n * m + 1))
matrix = [nums[r*n:(r+1)*n] for r in range(m)]
Amazing that this transposes a matrix:
list(zip(*matrix))
If I need to transpose in place (for a square matrix):
n = len(matrix)
for r in range(n):
for c in range(r + 1, n):
matrix[r][c], matrix[c][r] = matrix[c][r], matrix[r][c]
counts = {"a": 3, "b": 7, "c": 2}
max(counts, key=counts.get) # => "b"
If I need the max k, a counter is handy:
from collections import Counter
counts = Counter({"a": 3, "b": 7, "c": 2})
counts.most_common(k)
If I need a mutex around some shared state, use:
from threading import Lock
unfair_lock = Lock()
with unfair_lock:
# my critical section
To submit work to a pool of background threads, use:
from concurrent.futures import ThreadPoolExecutor
with ThreadPoolExecutor() as pool:
for res in pool.map(add_one, [1,2,3]):
print(
# Submit work
The API is very thin for ThreadPoolExecutor, just submit, map, and shutdown.
If I submit work with pool.submit, I get back a future. Work is scheduled immediately and does not block.
Gotchas:
- future.result blocks until that tasks’s result is ready
- pool.map(fn, argument) returns an in-order iterator of results. Calling next or iterating in a for blocks until that task’s result is ready.
- The main thread becomes the coordination thread, and it will block (!) until all work is done in the contextmanager, even if I don’t explicitly wait on results.
Example straight from the heap docs: https://docs.python.org/3/library/heapq.html#basic-examples
from heapq import heappush, heappop
h = []
heappush(h, (5, 'write code'))
heappush(h, (7, 'release product'))
heappush(h, (1, 'write spec'))
heappush(h, (3, 'create tests'))
heappop(h)
Another good one is heapify:
heap = [5, 2, 8]
heapq.heapify(heap)
heapq.heappush(heap, 1)
heapq.heappop(heap) # 1
heap[0] # 2
Ref:
| Operation | Meaning | Time |
|---|---|---|
heapq.heapify(nums) |
Turn a list into a heap, in place | O(n) |
heapq.heappush(heap, x) |
Add an item | O(log n) |
heapq.heappop(heap) |
Remove and return the smallest | O(log n) |
heap[0] |
Peek at the smallest | O(1) |
heapq.heappushpop(heap, x) |
Push, then pop the smallest | O(log n) |
heapq.heapreplace(heap, x) |
Pop the smallest, then push | O(log n) |
OrderedDict gives you move_to_end and popitem(last=False). Can use it for an easy LRU cache.
No need to manage a dict and doubly linked list myself.
Also, somewhere I got it in my head that I needed to evict before adding,
but I don’t. Add then evict makes the logic straightforward:
from collections import OrderedDict
class LRUCache:
def __init__(self, capacity):
self.capacity = capacity
self.od = OrderedDict()
def get(self, key):
if key not in self.od:
return -1
self.od.move_to_end(key)
return self.od[key]
def put(self, key, val):
self.od[key] = val
self.od.move_to_end(key)
if len(self.od) > self.capacity:
self.od.popitem(last=False)
Update: I know where evict before adding came from, that’s an LFU cache, because
if you insert first that k/v’s used counter is at one and would be immediately evicted.
Handy for creating hashes of arbitrary files:
import hashlib
hasher = hashlib.sha256()
with open(path, "rb") as f:
while chunk := f.read(64 * 1024):
hasher.update(chunk)
result = hasher.hexdigest()
Git uses SHA-1. There is also a direct form if I don’t need to stream in chunks:
hashlib.sha256(b'hello world')
If you forget to call to_thread on a sync function, you block the entire event loop!
It’s a big topic: https://docs.python.org/3/library/concurrency.html
Most used are the initializer with maxlen for built-in eviction, append, and popleft
and : deque(maxlen=3)
| Operation | Code | Returns / effect | Time |
|---|---|---|---|
| Initialize w eviction | deque(maxlen=3) |
new deque | O(1) |
| Initialize w existing | deque([1,2,3]) |
new deque | O(n) |
| Add right | q.append(4) |
[1, 2, 3, 4] |
O(1) |
| Add left | q.appendleft(0) |
[0, 1, 2, 3] |
O(1) |
| Remove right | q.pop() |
Returns 3 |
O(1) |
| Remove left | q.popleft() |
Returns 1 |
O(1) |
| Peek right | q[-1] |
3 |
O(1) |
| Peek left | q[0] |
1 |
O(1) |
| Length | len(q) |
3 |
O(1) |
| Empty check | not q |
False |
O(1) |
| Extend right | q.extend([4, 5]) |
[1, 2, 3, 4, 5] |
O(k) |
| Extend left | q.extendleft([4, 5]) |
[5, 4, 1, 2, 3] — reverses input |
O(k) |
| Rotate right | q.rotate(1) |
[3, 1, 2] |
O(k) for k steps |
| Rotate left | q.rotate(-1) |
[2, 3, 1] |
O(k) for k steps |
| Reverse in place | q.reverse() |
[3, 2, 1] |
O(n) |
| Membership | 2 in q |
True |
O(n) |
| Count | q.count(2) |
1 |
O(n) |
| Remove first match | q.remove(2) |
[1, 3]; raises ValueError if absent |
O(n) |
| Clear | q.clear() |
Empty deque | O(n) |
| Find index | q.index(val) |
Returns the first index of the value | O(n) |
| Insert at | q.insert(index, val) |
Inserts value before index | O(n) |